Long-Horizon Agents Need Experiments, Not Just Prompts - Erina Karati

Long-horizon multi-agent systems need an auto-research experimental loop - controlled scenarios, traced behavior, scorecards, and a small editable policy surfa...

By Sean Weldon

Long-Horizon Agents Need Experiments, Not Just Prompts

Abstract

Long-horizon multi-agent systems degrade in ways invisible to single-response evaluation: facts propagate but lose provenance, uncertain rumors harden into asserted truths, and agents hold knowledge without acting on it. This synthesis examines Project Paradox, a modular multi-agent framework built at Supercell's AI Innovation Lab, and the auto-research experimental loop layered atop it. The central thesis is that retrieval-augmented memory and manual prompt tuning are insufficient for sustained social consistency; systematic improvement requires controlled scenarios, structured behavioral traces, a balanced scorecard, and a deliberately constrained editable policy surface. The approach reframes evaluation from per-response quality to society-level behavior across an entire run. Findings indicate that small, measured protocol changes - preserving source attribution, encoding confidence markers, adjusting replanning triggers - produce disproportionate effects on emergent behavior, with implications extending to support, research, coding, and workflow agents.

1. Introduction

Autonomous agents built on large language models are increasingly deployed in settings where behavior must remain coherent over hundreds or thousands of interaction steps. Long-horizon behavior refers to an agent's ability to maintain consistent state, commitments, and knowledge attribution across extended timeframes - a property distinct from the per-turn fluency measured by most benchmarks. As agent deployments move from single-turn assistants to persistent, socially embedded systems, the gap between "sounds coherent in one exchange" and "remains coherent across a society over time" becomes the primary engineering obstacle.

Project Paradox is a modular framework enabling developers to embed intelligent autonomous agents into video games. Its agents move with intent guided by memory, emotion, or curiosity; interact with objects; remain context-aware of their environment and other characters; and react to events in ways that dynamically alter beliefs and emotional state. Agents may initiate conversations with other agents or human players, with those exchanges written to memory and subsequently shaping future emotions, beliefs, and goals.

The research question addressed here is how to systematically improve social consistency in such a system, rather than relying on memory retrieval alone or ad hoc prompt tuning. As stated directly in the source material: "Long horizon agents need experiments and not just prompts." Section 2 establishes the architectural foundation. Section 3 analyzes observed failure modes and the auto-research loop designed to address them. Section 4 consolidates actionable technical findings. Section 5 discusses generalization and limitations. Section 6 concludes.

2. Background and Related Work

The system's stateful architecture provides the substrate on which all later experimentation operates. Four mechanisms are central: a per-agent memory namespace backed by Retrieval-Augmented Generation (RAG), which prevents memory bleeding between agents; an emotion vector spanning joy, sadness, fear, anger, and disgust, updated after events and conversations; belief scores functioning as a trust matrix between agents and the player, adjusted by LLM-driven classification (up, down, unchanged); and an importance scoring mechanism that routes consequential memories into a separate high-fidelity retrieval cache.

The auto-research concept draws on Karpathy's formulation of automated experimentation. Critically, it is positioned as a meta system - external to the simulated world rather than an additional inhabitant of it. This distinction establishes the intellectual context for the analysis: evaluation infrastructure must sit outside the system under test, observing full traces rather than participating in them.

3. Core Analysis

3.1 The Long-Horizon Failure Mode

Short-horizon gameplay performed adequately: agents could plan, move, talk, and recall recent interactions. Over longer horizons, social consistency weakened in specific, reproducible ways. In one documented case, a rumor about a mango sale propagated between agents but lost its source attribution - the system retained the rough topic while discarding provenance. Separately, rumors could harden into stated facts rather than persisting as uncertain claims, and agents could possess a fact in memory yet fail to incorporate it into subsequent plans. These are not failures of retrieval per se; RAG successfully surfaced relevant memories. They are failures of what the retrieved content represents - confidence, source, and actionability were not preserved as first-class attributes.

3.2 The Auto-Research Loop

The auto-research layer addresses this by treating an entire simulation run, rather than a single agent response, as the unit of evaluation. The loop proceeds as: define a controlled scenario, run the simulation, collect structured traces, score behavior against a balanced scorecard, propose a constrained protocol change, rerun, and keep or revert the change based on measured improvement. This reframes the objective from "does this response look reasonable" to "does society-level behavior improve under measurement."

Three controlled scenarios illustrate the method: a public fact diffusion scenario testing whether the correct agents learn a fact and retain its source; a rumor uncertainty scenario testing whether "might leave" hardens into "is leaving"; and a replanning scenario testing whether agents update and communicate blocked routes. The source material notes that without such controlled scenarios, "wandering agents" make evaluating improvement very difficult - there is no ground truth against which drift can be measured.

3.3 Scorecard Design and the Editable Surface

A single vague metric such as "agent quality" hides specific failure modes. The balanced scorecard instead separates diffusion (reach: how many agents know a fact after N steps), provenance (source retention among agents who know the fact), rumor fidelity (uncertainty preservation and false certainty rate), planning (action consistency and time to replan), and privacy (containment). This decomposition matters because optimizing a single metric can induce bad behavior - for instance, maximizing diffusion reach can produce oversharing or leakage of information that should remain private.

Equally important is the constraint on what the auto-research layer may modify. The harness, scenarios, and metrics are frozen; only specific policy surfaces are exposed for editing: memory writing policy, retrieval policy, communication prompts, belief/trust rules, source attribution, and replanning triggers. This distinguishes a system that searches within a bounded policy space from one that permits an LLM to rewrite arbitrary code - a distinction the source material treats as essential to controllability.

4. Technical Insights

Several implementation-level findings emerge with direct applicability to agent system design:

5. Discussion

The broader implication is that long-horizon consistency problems are not unique to game-based agent simulations. The source material draws explicit parallels to support agents, which must track which policy update supersedes an older answer; personal assistants, which must remember and correct prior commitments; research agents, which require provenance citations and contradiction handling; coding agents, which must maintain context across issues, files, and changing requirements; and workflow agents, which require access control and replanning when conditions change. In each case, the underlying problem is the same: maintaining state that correctly informs future action, rather than merely being retrievable.

This suggests that current evaluation practice in the broader agent research community - heavily weighted toward single-turn or short-context benchmarks - under-measures the failure modes that matter most in deployed, persistent systems. A knowledge gap remains around how to generalize scenario design across domains without bespoke construction for each deployment, and how automated policy search scales as the editable surface grows beyond the six categories identified here.

6. Conclusion

This analysis demonstrates that sustaining social consistency in long-horizon multi-agent systems requires infrastructure beyond memory retrieval and manual prompting: frozen evaluation harnesses, controlled scenarios, structured traces, balanced scorecards, and a narrowly constrained editable policy surface. The practical takeaway is procedural - freeze the harness, define scenarios, log traces, score behavior, expose only a small policy surface, and retain only changes that survive measurement. This recipe generalizes beyond game agents to any system where an agent's past state must correctly shape its future action.


Sources


About the Author

Sean Weldon is an AI engineer and systems architect specializing in autonomous systems, agentic workflows, and applied machine learning. He builds production AI systems that automate complex business operations.

LinkedIn | Website | GitHub