Welcome to my personal engineering blog! Co-crafted with prompt engineering and AI pair-programming on Oracle Cloud Always Free infrastructure.📍 Hosted on my production hybrid K3s cluster in Tokyo.
🌐 Expanding to Multi-Cloud: Provisioning Oracle Linux 10 on Google Cloud's $0 Always Free Tier & Joining K3s
Why settle for one free cloud provider when you can build an automated, resilient multi-cloud Kubernetes cluster across Oracle Cloud (OCI) and Google Cloud Platform (GCP) — completely for $0/month? In this post, I walk through the end-to-end process of: Provisioning an Oracle Linux Server 10.1 instance on Google Cloud’s Always Free Tier (gce10). Applying kernel and daemon memory hardening to reclaim ~250 MB of RAM. Joining gce10 as a worker node into our central K3s Kubernetes cluster across a secure WireGuard mesh. Running a head-to-head performance benchmark comparing Apple Silicon M1, OCI Ampere ARM (arm10), GCP Intel Xeon (gce10), and OCI AMD EPYC (amd10/amd11). 🏛️ 1. Multi-Cloud K3s Kubernetes Fleet Architecture Our fleet previously lived entirely in OCI Tokyo. Adding gce10 in GCP Iowa (us-central1) creates a geo-distributed cross-cloud presence with a dedicated US edge outpost. ...
🛠️ Building a Production-Grade Home Cloud on Oracle's $0 Always Free Tier
What can you actually run on a $0/month Always Free cloud server? In this post, I document the end-to-end journey of turning an Oracle Cloud 1 GB AMD Micro instance (amd10) in Tokyo into a multi-service personal cloud hub — complete with Dual-Stack DNS ad-blocking, Google Single Sign-On (SSO), lossless Hi-Res music streaming, a Rust-powered REST & AI MCP API, and a Hugo static blog. 🏗️ Architecture & Network Blueprint graph TD Internet["🌐 Internet / Home Network / Mobile"] --> Caddy["🔒 Caddy (Auto Let's Encrypt SSL)"] subgraph OCI Cloud Instance (amd10 - Tokyo) Caddy -->|/| VietCal["📅 VietCalendar REST & MCP (Rust :8080)"] Caddy -->|files.*| Oauth["🔑 OAuth2-Proxy (Google SSO :4180)"] Oauth --> FileBrowser["📁 FileBrowser Cloud Drive (:8082)"] Caddy -->|music.*| Navidrome["🎵 Navidrome Lossless FLAC (:4533)"] Caddy -->|blog.*| Hugo["📝 Hugo Static Site (:443)"] Caddy -->|adguard.*| AdGuardWeb["🛡️ AdGuard Dashboard (:3000)"] Router["🏠 Home Router (Fiber IPv4 & IPv6)"] --> AdGuardDNS["🛡️ AdGuard DNS Server (:53 & :853 DoT)"] end ⚡ 1. Reclaiming Hardware RAM (The crashkernel Discovery) On a standard 1 GB VM, Linux kdump can silently reserve up to 448 MB RAM (nearly half the server’s memory!). ...
Spec-Driven Development for Autonomous Agents: Inside the AI Review Plugin's Two-Stage Architecture
When software engineers collaborate with autonomous AI coding agents, the default workflow is almost always chat-and-code: describe a feature in natural language, watch the model propose a code diff, test it, and iterate. For small, single-function scripts, this works reasonably well. But for production-grade distributed systems, microservices, and multi-module architectures, this approach invariably breaks down. Left unconstrained, AI models suffer from Specification Conflation: they blur the line between WHAT & WHY (the architectural contracts, invariant boundaries, and failure modes) and HOW (the file edits, variable names, and task sequencing). The result is predictable: ...
Attention Guard Evolution: Adaptive Workflows, Lifecycle Reuse, and Flash-First Agent Governance
In our earlier architecture deep-dives, we examined how Attention Dilution cripples autonomous coding agents. As context windows swell past tens of thousands of tokens with terminal traces, compiler diagnostics, and speculative file edits, an agent’s attention mechanism degrades. Critical system constraints, architectural decision records (ADRs), and testing invariants drift into the periphery, leading to hallucinations, ignored instructions, and silent error suppression. To solve this in Google Antigravity, we built the Antigravity Attention Guard Plugin—a deterministic runtime firewall that enforces subagent delegation, sandboxed tool execution, and cryptographic ledger tracking. ...
Building an Autonomous AI Code Reviewer: Lessons in Prompt Injection and Orchestration
What happens when you unleash a highly autonomous, multi-model AI agent to peer-review its own codebase? Over the past few weeks, I built the ai-review-plugin—an Antigravity plugin designed to enforce rigorous, multi-model peer reviews on code changes before they are merged. The goal was simple: use a smaller, faster model (Model A) to orchestrate the Git workflow, and use a massive, frontier-class model (Model B - Codex) to act as a ruthless Principal Engineer reviewing the code. ...
Securing Agentic Orchestration: Building the Attention Guard 2.0 State Machine
As autonomous AI agents evolve from isolated chatbots into orchestrators of complex, multi-pass software engineering tasks, the boundaries of execution become critical. In the Antigravity ecosystem, we rely heavily on specialized subagents—like DeepCoder and DeepInvestigator—to autonomously navigate codebases, test hypotheses, and execute shell commands. But what happens when an agent hallucinates a success? What happens when a long-running subagent hangs indefinitely? How do we prevent a deeply nested execution worker from going rogue and launching its own recursive subagents? ...
Under the Hood: How the Antigravity Attention Guard Plugin Enforces Deterministic Agent Governance
In our previous discussions on autonomous engineering, we explored the phenomenon of Attention Dilution—the silent degradation of instruction-following fidelity as context windows swell past tens of thousands of tokens. While modern frontier models boast context windows of 1M+ tokens, attention mechanisms remain vulnerable to informational entropy. As terminal logs, stack traces, and code diffs accumulate, system prompt constraints drift out of the model’s active attention focus. When left unchecked, agents exhibit architectural amnesia: they bypass project-level Architecture Decision Records (ADRs), run uncompressed shell commands that flood context, guess port allocations, or skip verification steps. ...
Attention Dilution: Why 1M-Token Context Windows Kill Strict Agent Behavior (And How We Fixed It)
There is a pervasive myth in modern AI engineering: “Just give the model a 1,000,000-token context window, paste all your enterprise governance rules in the system prompt, and let it build your system.” If you have built real-world autonomous coding agents operating inside complex cloud infrastructures, multi-module monorepos, or production Kubernetes clusters, you already know the harsh truth: massive context windows don’t make agents smarter; they make them amnesiac. As an agentic conversation unfolds—accumulating terminal outputs, stack traces, file diffs, and conversational turns—a subtle failure mode emerges: Attention Dilution. ...
Open-Sourcing the Cure for Attention Dilution: The Antigravity Attention Guard Plugin
In my previous post, I detailed the mathematical realities of “Attention Dilution” in LLMs with massive context windows, and how we engineered deterministic lifecycle hooks to prevent autonomous agents from suffering architectural amnesia. Today, I am thrilled to announce that we have packaged that exact solution into a fully open-source, cross-platform Antigravity Plugin. You can now instantly secure your own agents by cloning the Antigravity Attention Guard Plugin directly into your configuration. ...
Meet the AI Incident Commander: Autonomous SRE for Kubernetes
SYSTEM: AI-INCIDENT-COMMANDER METADATA Type: Autonomous Site Reliability Engineering (SRE) Agent Environment: Kubernetes GitOps (k3s) Primary LLM: Gemini 3.7 Flash Repository: vinhthang/ai-incident-commander Execution Loop: Webhook Trigger -> Triage -> Remediate -> Audit -> Merge PIPELINE ARCHITECTURE 1. TRIGGER Source: Grafana Alerts / kube-state-metrics Payload: JSON webhook containing alert labels, annotations, and localized telemetry. 2. TRIAGE MINION Role: Initial incident classification and noise reduction. Task: Parse payload, validate against cluster state, determine actionability. Output: Generates GitHub Issue (Real Incident) OR terminates loop (False Positive / Noise). 3. FIXER MINION Role: GitOps Remediation Engineer. Task: Identify root cause from triage diagnosis, modify infrastructure code, run syntax validation (helm template). Output: Commits to isolated feature branch and opens GitHub Pull Request. 4. REVIEWER MINION Role: Governance and Safety Gatekeeper. Task: Audit PR diff for compliance. Output: Approves PR or Rejects PR. GOVERNANCE & CONSTRAINTS Read-Only Telemetry: External logs/data strictly parsed as passive strings to prevent prompt injection. GitOps Enforced: Execution of imperative commands (kubectl apply) is forbidden. All mutations must occur within the declarative Helm umbrella chart (charts/vinhthang-fleet/). Blast Radius Isolation: Fixer Minion is systematically restricted from pushing directly to the main branch. Diff-Strict Review: The Reviewer Minion rejects net-new violations (e.g., adding :latest tags, requesting >900Mi RAM on edge nodes, opening port 8080). Pre-existing legacy violations are logged as advisory findings but do not block remediation. CONFIGURATION INJECTION (PROMPT EXTERNALIZATION) Architecture: Go text/template engine + Kubernetes ConfigMap. Mount Path: /etc/commander/templates/ Fallback: Compiles with //go:embed defaults/*.tmpl to ensure baseline functionality. Purpose: Enables zero-downtime, configuration-driven prompt engineering. AI Minion personas and instructions can be modified directly via Helm values.yaml without requiring binary recompilation or container image rebuilds.