How We Built an Agent That Improves Itself - Zubin Aysola, Weights & Biases
Weights and Biases built the Arya agent using a tight production-to-offline evaluation flywheel where production traces are continuously mined to create offl...
By Sean WeldonHow We Built an Agent That Improves Itself: A Synthesis of the Weights & Biases Arya Architecture
Abstract
This synthesis examines an engineering methodology for evaluating and improving production AI agents, drawn from the development of Arya, an agent built by Weights & Biases (W&B). The central claim is that benchmarks, evaluation harnesses, and agent configurations are mutually covariant, rendering static evaluation inadequate for dynamically evolving systems. The proposed remedy is a continuous production-to-offline flywheel in which live production traces are mined, replicated into sandboxed simulation environments, and converted into offline hill-climbing tasks. Byte-identical agent code executes in both settings, with a four-hour synchronization job preventing drift. The approach yielded 886 categorized offline tasks, three complementary task-simulation modalities, and nightly continuous integration (CI) runs over seven weeks reaching approximately 66% performance on certain task classes. Notably, the agent was demonstrated executing this research loop on itself, including autonomous bug identification and repair, suggesting a viable path toward self-reinforcing agent development pipelines.
1. Introduction
Agentic systems built atop large language models (LLMs) resist conventional evaluation practice. A static benchmark presumes a fixed system under test; agents, by contrast, change continuously as prompts, tool definitions, context-management policies, and underlying models are revised. This difficulty is captured directly: "Benchmarks, evaluations, the agents, and how you configure them are all covariant." Because every element of the measurement apparatus co-evolves with the artifact being measured, a score obtained today may not be commensurable with one obtained a week later.
This covariance problem is not merely theoretical. Teams building production agents routinely observe that improvements validated offline fail to materialize in deployment, or conversely, that production failures cannot be reproduced in test harnesses. The W&B engineering team's response was architectural rather than statistical: rather than attempting to hold the benchmark fixed, they built infrastructure that allows the benchmark itself to evolve in lockstep with production reality.
The central thesis of this work is that a production-to-offline evaluation flywheel - continuous mining of production traces to generate offline hill-climbing tasks - provides a tractable basis for principled agent improvement, and that a sufficiently capable agent can operate this loop largely autonomously. This analysis covers the observability substrate enabling the flywheel, the synchronization architecture preventing drift, the harness and sandbox design philosophy, the scoring methodology, and a live demonstration in which the agent under study, Arya, was shown identifying and repairing a bug in its own evaluation infrastructure.
2. Background and Related Work
The problem is framed by explicit analogy to the sim-to-real gap in reinforcement learning (RL): a policy optimized in simulation degrades when transferred to a physical environment because the simulator omits or distorts relevant dynamics. For agentic software, the offline evaluation harness plays the role of simulator, and the deployed product constitutes the real environment. Improvements measured offline may not transfer if the offline task distribution diverges from the production distribution.
Two RL-derived heuristics further inform the design. First, the adage that "it's just better to run more experiments than fewer" motivates cheap, parallel instantiation of agent variants via YAML-defined configurations. Second, the observation that contemporary frontier models already perform competently at W&B-domain tasks motivates a deliberate preference for prompt engineering and skill construction over reinforcement learning-based policy training. The internal designation for the offline benchmarking principle is the Weights & Biases Agent Factory (WBAF). The observability substrate underpinning the entire architecture is Weave, W&B's tracing platform, used symmetrically for offline simulation and production deployment.
3. Core Analysis
3.1 Symmetric Observability as the Enabling Substrate
The architecture's foundational commitment is that production and offline traces share an identical logging format within Weave. On the offline side, simulation environments are constructed and Arya is run within them, with results tracked in Weave. On the production side, the deployed agent logs in the same format, which permits production traces to be pulled directly into offline environments for replay. This symmetry is what makes hill-climbing on production errors and offline metrics possible within a single unified system, rather than requiring separate, incommensurable measurement pipelines.
3.2 Production-Research Synchronization
A critical architectural decision is that research and production code are byte-wise identical - the same agent is benchmarked in both environments. This eliminates a common failure mode in which an offline-optimized variant diverges silently from what is actually deployed. To maintain this identity over time, a four-hour sync job runs between the production and research environments, preventing code drift. Evaluation runs generate score trajectories against multiple logged metrics, providing a continuous record of how candidate variants compare to the production baseline.
3.3 Harness Design and the Prompt-Engineering-over-RL Decision
The harness design philosophy explicitly favors prompt engineering over reinforcement learning training, on the grounds that sophisticated foundation models already perform well at W&B-specific tasks without additional policy optimization. Engineering effort is instead directed toward building skills - modular capabilities - for the software agent, alongside an agnostic software stack handling compaction, context preparation, and UI payload assembly. YAML-defined configurations allow rapid creation of multiple parallel agent variants, operationalizing the "more experiments than fewer" heuristic at low marginal cost.
3.4 Sandbox Architecture and Emergent Behavior
Simulation environments follow an agnostic DAG pattern: configuration (YAML) → hydrate → set up environment → rehydrate (hot-patch runtime configuration) → run agent → score. These environments are computationally expensive, involving full machine learning training logs and simulated GPU executions, and are torn down after parallel runs to avoid interfering with concurrent work. Notably, the sandbox is described as unconstrained, permitting emergent agent behavior; in one instance, Arya was observed constructing its own parallel execution environment within the sandbox, an outcome not explicitly engineered by the development team.
3.5 Scoring Methodology and Task Corpus
Two scoring patterns are employed: normative scoring, which evaluates simple pass/fail outcomes against a task, and relativistic scoring, which compares variant behaviors against one another (for example, whether a variant asks clarifying questions versus proceeding without them). A team member conducted a dedicated audit of evaluation health, examining drift between offline evaluations and production behavior. The resulting corpus comprises 886 tasks, categorized by difficulty level and exposed to the product team for relevance review. Tasks are simulated through three complementary methods: simple text instructions, persona-based multi-turn user simulation, and direct replication of production traces - the latter providing the tightest coupling between the offline harness and real-world usage patterns.
4. Technical Insights
Several implementation-relevant findings emerge from this architecture. The 886-task corpus, organized by difficulty level, functions as a hill-climbing target set that is explicitly kept current through product-team review rather than being frozen at creation time. The four-hour synchronization cadence represents a specific, tunable trade-off between drift risk and system overhead. Nightly CI jobs tracking production and candidate agent variants over a seven-week period achieved approximately 66% performance on certain task classes, providing a longitudinal signal of both agent capability and evaluation stability. Tasks are specified as YAML documents defining starting conditions, user configurations, and ending conditions, which standardizes task authoring across the three simulation modalities. In one documented instance, Arya identified a bug involving an improper weave.log SDK call within the sandbox, authored a corresponding hill-climb target, and resolved the issue through a prompt injection into an existing skill - illustrating a closed-loop bug discovery and repair cycle. Separately, the CI pipeline itself broke and was reported to have been self-fixed by Arya overnight, indicating that the self-improvement loop extends to infrastructure maintenance, not solely task performance.
5. Discussion
These findings bear on a broader industry question: how should organizations evaluate agents whose behavior, tooling, and underlying models change on a near-continuous basis? The W&B approach suggests that the answer is not a better static benchmark but an evaluation architecture that treats the benchmark as a live artifact, continuously re-derived from production. This reframes evaluation engineering as a first-class, ongoing discipline rather than a one-time gate.
The demonstration of Arya operating this loop on itself - reviewing its own codebase, launching training jobs, mining production traces for new hill-climb tasks, and repairing its own bugs - raises questions about the appropriate boundary of agent autonomy in research infrastructure. While the quoted claim that a team member "haven't written a line of code in maybe eight months" is anecdotal, it signals a trajectory in which agent-authored evaluation and remediation become normalized practice. Open questions remain regarding failure detection when the agent's self-assessment is itself flawed, and regarding the long-term stability of a corpus whose tasks are generated by the same system being evaluated.
The explicit sim-to-real framing also suggests that lessons from robotics and RL regarding environment fidelity, domain randomization, and transfer validation may generalize productively to LLM agent evaluation, an area where such analogies remain underexplored in practice.
6. Conclusion
This synthesis has described a production-to-offline evaluation flywheel built around byte-identical code paths, symmetric observability, a large categorized task corpus, and a sandbox architecture permitting emergent agent behavior. The core contribution is architectural: rather than resolving the covariance between agents, evaluations, and configurations through fixed benchmarks, the system embraces covariance by continuously regenerating its evaluation targets from production reality. The practical takeaway for engineering teams is that investment in symmetric logging infrastructure and drift-preventing synchronization may yield more durable evaluation practice than investment in any single static benchmark, with the replication pattern from production to simulation to defined improvement representing the architecture's central, reusable insight.
Sources
- How We Built an Agent That Improves Itself - Zubin Aysola, Weights & Biases - Original Creator (YouTube)
- Analysis and summary by Sean Weldon using AI-assisted research tools
About the Author
Sean Weldon is an AI engineer and systems architect specializing in autonomous systems, agentic workflows, and applied machine learning. He builds production AI systems that automate complex business operations.