Get Out of the Model's Way - Kevin Hou, Google Antigravity
Product design must scale with model intelligence - getting out of the model's way with new primitives like sub agents, sidecars, and generative UI defines the...
By Sean WeldonGet Out of the Model's Way: Primitives for the 2026 Agentic Coding Era
Abstract
This synthesis examines a product-design thesis for agentic coding systems: as frontier model capability increases, product architecture must evolve to expose that capability rather than constrain it. The analysis draws on the development trajectory of Google's Antigravity, an agent-first integrated development environment (IDE) launched in November 2025, and its successor release, which decoupled the IDE from an agent manager into two distinct applications. Three primitives are identified as characteristic of the 2026 agentic era: dynamically generated sub agents, sidecars (long-lived trigger-driven processes), and generative UI. Empirical evidence includes a hero run in which 93 sub agents produced an operating-system kernel over 12 hours using 15,000 requests and two billion tokens for under $1,000, and an internal research-evaluation workflow reduced from manual notebook analysis to a largely automated minutes-long pipeline. Implications for orchestration design, cost modeling, and interface generation are discussed.
1. Introduction
The design of developer tooling built atop Large Language Models (LLMs) has, until recently, assumed that the model is a component embedded within a fixed product surface. The central thesis examined here inverts that assumption: product design must scale with intelligence, meaning that each capability jump in the underlying model should be matched by a corresponding expansion of the product's primitives. As articulated in the source material, "LLMs aren't just role players anymore. They can be your star player if you build the right product around them," and consequently, "to let your star player cook, you have to get out of the model's way."
Two terms require early definition. An agent manager is an orchestration surface for launching, monitoring, and coordinating many concurrent agents, distinct from an IDE, which provides fine-grained inspection of code and execution state. A primitive is a composable, model-facing abstraction - such as a tool, permission scope, or process type - that defines what an agent can do within a product.
This paper analyzes the primitive-by-era evolution of agentic coding products, the organizational friction accompanying each transition, the three primitives proposed for the 2026 era - sub agents, sidecars, and generative UI - and the quantitative evidence supporting their viability, drawn from Google's Antigravity product line and internal research automation examples.
2. Background and Related Work
The source material organizes recent history into three capability eras, each defined by the primitives it unlocked. The 2022 era of autocomplete and chat sidebars relied on embeddings, rules files, and syntax-tree parsing, with fully deterministic control flow invoking the model at fixed points in a human-authored pipeline. The 2024 era introduced agents capable of autonomous multi-step execution, requiring new primitives: Model Context Protocol (MCP) servers, custom tools, and permission systems to bound agent authority. The 2025 era of agent managers, running many agents in parallel, introduced skills, hooks, and artifacts as the coordination substrate.
Antigravity itself illustrates this progression. The initial release positioned an agent-first IDE around the agent manager concept for technical and non-technical users. A command-line interface was subsequently extracted and shipped standalone. The 2.0 release, announced at Google I/O, formally separated the IDE and agent manager into two applications and added sub agents, new models, git worktrees, scheduled tasks, and voice mode. A second relevant framework is the research-product flywheel: because Google DeepMind controls both model training and the product surface, observed product usage can feed back into training objectives, tightening the iteration loop between capability and interface.
3. Core Analysis
3.1 Transition Costs and User Resistance
Each primitive shift incurred measurable adoption friction. Granting agents terminal access provoked fears of catastrophic, irreversible actions - referenced in the source material via the "son of Anton" incident involving codebase deletion. This friction was resolved not by withdrawing terminal access but by pairing it with permission systems and improving model reliability, which ultimately allowed users to ship faster despite initial safety concerns. A second instance of resistance occurred when the chat sidebar was removed from Windsurf, generating user backlash; this was eventually superseded as multistep agentic research and execution replaced the chat sidebar as the preferred interaction paradigm. These episodes suggest that primitive transitions are not frictionless even when they represent net capability improvements, and that user trust must be rebuilt at each architectural boundary.
3.2 IDE-Agent Manager Decoupling
The decoupling of the IDE from the agent manager is framed through an explicit analogy: "the IDE is to the agent manager what the debugger was to the IDE" - a tool not always necessary but valuable for deeper inspection when required. This reframing positions the agent manager, not the IDE, as the primary daily surface, with the IDE reserved for cases demanding granular code-level control. The source material predicts that agent orchestration - agent teams, swarms, or "software factories" - constitutes the forward trajectory for the category, enabled in part by Google DeepMind's tight feedback loop between product telemetry and model training.
3.3 The 2026 Primitive Set: Sub Agents, Sidecars, Generative UI
Three primitives are proposed as defining the 2026 agent-team era. Gemini 3.5 Flash, launched in April, was optimized specifically for leading teams of agents rather than merely executing discrete tasks, and underpins a /teamwork slash command that launches swarm mode with a lead agent managing an arbitrary-sized team.
Sub agents are dynamically generated and specialized - spanning frontend, backend, infrastructure, QA, and design roles - and may run a different underlying model than the main orchestrating agent. Demonstrated applications include an in-browser raw photo editor, a messaging application, and a full operating-system kernel built from scratch, capable of running Doom. The kernel build consumed 93 sub agents over 12 hours, issuing 15,000 requests and consuming 2 billion tokens, for a total cost under $1,000 - evidence offered in support of the scalability and economic feasibility of large-scale sub agent orchestration.
Sidecars are long-lived utility processes that listen for external triggers such as SMS messages, webhooks, cron jobs, or GitHub pull requests. This primitive already powers scheduled tasks within Antigravity, and a formal specification is slated for release in the coming summer.
Generative UI is enabled by the inference speed of Gemini 3.5 Flash, which renders at approximately 900 tokens per second - roughly 10x faster than other frontier model experiences. This throughput makes inline, dynamically generated interfaces viable as a replacement for static templates or hand-authored HTML, supporting the claim that "human written specialized UIs are dead."
3.4 Internal Research Automation as a Case Study
An internal side-by-side evaluation workflow illustrates the composite application of these primitives. A natural-language request triggers an agent that understands Google's monorepo via skills, computes performance deltas, and hands off to a research agent specialist that proposes approximately 100 hypotheses explaining the observed deltas. Sub agents are then spun up per hypothesis to investigate in parallel, with results mapped back into a single consolidated report. A generative UI is produced at the end, allowing interactive filtering and segmentation of findings. This workflow, previously requiring manual Jupyter notebook analysis, is now approximately 90% automated and completes in minutes.
4. Technical Insights
Several implementation-relevant findings emerge from the source material. First, sub agent heterogeneity - including the ability to select different underlying models per sub agent - suggests an orchestration architecture in which task-model fit, rather than a single fixed model, governs cost and quality trade-offs. Second, the kernel-build metrics (93 sub agents, 15,000 requests, 2 billion tokens, sub-$1,000 cost, 12 hours) indicate that large-scale parallel orchestration is economically viable at current token pricing, though the generalizability of this cost profile to less structured tasks remains untested in the source material. Third, the sidecar primitive's reliance on external triggers implies a shift toward event-driven agent activation rather than purely interactive, user-initiated sessions, which has implications for infrastructure design around persistent process management and trigger reliability. Fourth, the Gemini 3.5 Flash throughput figure (~900 tokens/sec) is presented as a necessary precondition for generative UI, indicating that interface-generation strategies are directly bottlenecked by inference speed rather than model capability alone.
5. Discussion
The broader implication of this thesis is that product architecture in agentic coding is not a fixed scaffold but a continually renegotiated boundary that must be redrawn as model capability advances. The recurring friction documented across transitions - terminal access, sidebar removal - suggests that user resistance to primitive changes is a predictable cost of this renegotiation, not evidence that the change was misguided. This has implications for how product teams communicate and sequence primitive rollouts.
A notable gap in the source material is the absence of failure-mode data: the kernel-build and research-automation examples are presented as successes, with no discussion of error rates, hallucinated hypotheses, or sub agent coordination failures at scale. Future investigation should address how orchestration quality degrades as sub agent counts increase, and what verification mechanisms are necessary when generative UI and sub agent outputs are consumed with minimal human review.
The research-product flywheel framework also raises a structural question for organizations without integrated model-training capability: whether the primitives described here are replicable by product teams lacking direct influence over model optimization, or whether they are contingent on the specific co-design relationship between Google DeepMind's training and product teams.
6. Conclusion
This analysis has traced a three-era progression in agentic coding products - from deterministic autocomplete, to agents with bounded permissions, to agent managers - and argued that the 2026 era is defined by three further primitives: sub agents, sidecars, and generative UI. The kernel-build and research-automation case studies provide quantitative grounding for the claim that large-scale, dynamically orchestrated agent teams are both technically feasible and economically modest in cost. Practically, teams building agentic products should treat primitive expansion as a recurring design obligation tied to model capability jumps, anticipate user resistance as a structural cost rather than a signal of failure, and evaluate inference speed as a direct constraint on interface-generation strategies such as generative UI.
Sources
- Get Out of the Model's Way - Kevin Hou, Google Antigravity - Original Creator (YouTube)
- Analysis and summary by Sean Weldon using AI-assisted research tools
About the Author
Sean Weldon is an AI engineer and systems architect specializing in autonomous systems, agentic workflows, and applied machine learning. He builds production AI systems that automate complex business operations.