Building Self-Improving Agent Software Factories - Suraj Gupta, Warp

Software factories themselves can self-improve over time through three concrete mechanisms - skills, persistent memory, and model routing - enabling agents to ge...

By Sean Weldon

Building Self-Improving Agent Software Factories

Abstract

Autonomous coding agents deployed at scale - "software factories" - typically require continuous human curation of prompts, procedures, and model assignments. This synthesis examines a three-mechanism framework for making such systems self-improving: skills functioning as procedural memory, persistent memory functioning as a scoped fact store, and model routing functioning as a cost-capability allocator. Drawing on implementation experience from Warp's agentic development environment (approximately one million active users) and its Oz cloud agent platform, the analysis describes an outer-loop/inner-loop architecture in which a supervisory agent observes execution trajectories and feedback signals to refine procedural artifacts, with updates version-controlled through git and gated by human pull-request review. Findings indicate that persistent memory reduces redundant context-gathering token expenditure, and that best-of-k evaluation across model candidates enables Pareto-efficient task-to-model assignment. Model routing remains largely heuristic, representing the principal open problem.

1. Introduction

The deployment of large language model (LLM) agents into production software engineering workflows has shifted the bottleneck from agent capability to agent maintenance. A software factory - an orchestrated collection of agents performing recurring engineering tasks such as issue triage, continuous integration (CI) repair, root-cause analysis, and documentation generation - accumulates configuration debt over time. Procedures written for one codebase state become inaccurate; context painstakingly assembled during one run is discarded and re-derived in the next; and model assignments fixed at design time become economically or qualitatively suboptimal as the model landscape evolves.

The central thesis examined here is that software factories can improve themselves along these three axes without proportional human effort. As stated directly in the source material: "Self-improvement is about agents being able to improve over time... phasing out humans from that loop, so that the improvement loop can be automatic."

Three terms require definition upfront. Skills are declarative descriptions of procedures an agent follows to complete a recurring task class - for example, a triage agent's method for reproducing a reported issue. Persistent memory is a versioned store of facts, learnings, and outcomes scoped to a particular agent, distinct from procedural knowledge. Model routing is the rule-governed mapping of task classes to specific underlying models, intended to balance capability against cost. This analysis proceeds by situating these constructs against existing agent-design patterns, then examines each mechanism in turn, extracts implementation-level findings, and concludes with implications for practitioners building similar systems.

2. Background and Related Work

The framework rests on a distinction familiar from cognitive psychology between procedural memory (knowing how) and declarative memory (knowing that). In agentic systems this distinction maps cleanly onto two separable artifacts: skill definitions encoding workflow steps, and memory records encoding environment-specific facts. Conflating the two produces artifacts that are simultaneously too rigid to generalize across contexts and too specific to reuse across tasks.

A second foundational pattern is the outer-loop/inner-loop agent architecture. The inner loop is the task-executing agent; the outer loop is a supervisory agent whose input is the inner loop's trajectory, outcome, and associated human feedback, and whose output is a modification to the inner loop's governing artifacts. This generalizes reflection and self-critique patterns by moving critique off the critical path of task execution and persisting its output as a durable, inspectable artifact rather than transient context. A third relevant precedent is treating model selection as an optimization problem over a Pareto frontier of cost against capability, which Warp's auto models feature operationalizes at product scale.

3. Core Analysis

3.1 Skills as a Self-Correcting Procedural Layer

Skills degrade predictably: "Skills get stale over time. Humans give feedback to the agents, agents learn things in their own runs and trajectories, and the skill that you previously gave it becomes dated." Warp addresses this through an outer loop agent that continuously observes inner loop runs of a given skill - such as the triage agent shipped in Warp's open-sourced client repository - and synthesizes updates from feedback signals including thumbs up/down ratings and free-text comments.

Critically, the mechanism preserves human oversight without requiring human initiation. Skill updates are committed via git and surfaced as pull requests rather than applied silently. This yields, in the source material's words, "full observability into how that skill is transformed over time," while ensuring "a human is actually going to review those updates to the skill" before changes take effect. The design choice separates detection of drift (automated) from authorization of change (human-gated), which limits the risk of runaway or degenerate procedural drift while still removing the burden of manual skill authorship.

3.2 Persistent Memory as a Fact Store

Where skills encode procedure, persistent memory encodes outcome-specific facts scoped to an individual agent. The motivating failure mode is redundant computation: "What happens when your Sentry agent looks for an issue and finds an issue... but then you encounter a similar issue in the future... you've probably wasted a lot of tokens re-gathering context." Persistent memory eliminates this redundancy by allowing an agent to retrieve prior learnings rather than re-derive them.

In Oz, Warp's cloud agent platform, memory stores are attachable to specific agents and are versioned, with each memory traceable to the run that produced it. Humans retain create/update/delete control, but retrieval and application during subsequent runs is automatic - illustrated by a Sentry-integrated agent drawing on five relevant memories during root-cause analysis without explicit human prompting. Notably, this memory layer operates uniformly "across all harnesses (Warp Zone, proprietary harness, plot code, Kodak)," indicating that the fact store is architected as an infrastructure-level capability rather than a per-integration feature.

3.3 Model Routing as Cost-Capability Allocation

The third mechanism addresses economic sustainability. Running every agent invocation - including low-complexity tasks such as triage or CI fixes - on a high-capability model like Opus is, in the source material's framing, "prohibitively expensive." Warp's response operates on two tiers. First, out-of-the-box auto models are continuously re-evaluated for Pareto efficiency as new models are released, removing the need for users to manually track the model landscape. Second, users can define custom rule-based routing configurations mapping specific task classes to specific models - for instance, routing database migration tasks to GLM and runbook or API documentation tasks to Quen.

The source material is explicit that this remains an underdeveloped discipline: "Model routing decisions are currently more art than science." Warp's internal approach to formalizing it is a best-of-k evaluation sidecar, which runs multiple candidate models against the same prompt to empirically determine optimal task-model pairings. One reported finding from this process is that UI-oriented tasks perform adequately on GLM, reducing dependence on more expensive models for that task class.

4. Technical Insights

Several implementation patterns generalize beyond Warp's specific stack. The outer loop/inner loop separation is architecturally significant because it decouples task execution latency from improvement latency - the inner loop need not slow down to incorporate learning, since refinement happens asynchronously and is reviewed before redeployment. This is reinforced by the git-based, pull-request-gated update mechanism, which provides an audit trail and a natural rollback point, addressing a common objection to autonomous self-modification: loss of control.

The versioning and provenance tracing applied to memory stores is a second transferable pattern - every stored fact is traceable to its originating run, which supports debugging when an agent's downstream behavior is affected by a specific historical memory. A trade-off worth noting is that automatic leveraging of past memories in future runs, while efficient, introduces a dependency on memory quality; erroneous or outdated facts could propagate silently unless the human-facing create/update/delete controls are actively exercised.

The best-of-k evaluation methodology for model routing is notable for treating model assignment empirically rather than heuristically, though the source material concedes the practice has not yet been formalized into customer-facing evaluation tooling. This represents a limitation: routing rules (e.g., GLM for database migrations) currently reflect internal findings rather than a generalizable, continuously validated framework.

5. Discussion

Taken together, the three mechanisms suggest a broader design principle: self-improvement is achieved not by increasing agent autonomy at the point of action, but by inserting supervisory and memory layers around execution that operate on longer time horizons and are subject to lighter-weight human gating than the original task itself. This reframes "phasing out humans from the loop" not as removing human judgment, but as relocating it from high-frequency, low-leverage decisions (writing every prompt or reviewing every model choice) to lower-frequency, high-leverage decisions (approving a skill update, curating a memory store).

An open question is how model routing evolves from art to science. The reliance on best-of-k sampling is empirically sound but resource-intensive, and its outputs (e.g., task-to-model mappings) are currently static rules rather than continuously adaptive policies. Given that Warp explicitly frames formal customer-facing evals as future work, this remains the least mature of the three mechanisms.

A second gap concerns interaction effects between mechanisms. The source material treats skills, memory, and routing as largely independent, but a skill update driven by feedback could plausibly interact with which model is best suited to execute it - an area not addressed in the available material.

6. Conclusion

This analysis has examined three concrete mechanisms - skills, persistent memory, and model routing - by which agentic software factories can reduce their dependence on continuous human maintenance. Skills function as version-controlled, human-reviewed procedural memory refined by an outer loop agent; persistent memory functions as a traceable, versioned fact store that prevents redundant context-gathering; and model routing functions as an empirically-driven, if still heuristic, cost-capability allocator.

The practical takeaway for practitioners building similar systems is that self-improvement need not mean full autonomy: git-gated skill updates and human-controlled memory stores demonstrate that automation and oversight are compatible design choices. Future work, per the source material, should prioritize formalizing model routing evaluation into a systematic, customer-facing framework rather than relying on internal best-of-k experimentation alone.


Sources


About the Author

Sean Weldon is an AI engineer and systems architect specializing in autonomous systems, agentic workflows, and applied machine learning. He builds production AI systems that automate complex business operations.

LinkedIn | Website | GitHub