No, That's Not a Software Factory - Ryan Cooke, WorkOS
Software factories should be measured by outcome metrics (feature delivery speed, customer impact) rather than output metrics (number of PRs or lines of code...
By Sean WeldonNo, That's Not a Software Factory: Reframing Outcome Metrics in Agentic Engineering Systems
Abstract
This synthesis examines WorkOS's production deployment of an agentic software factory and argues that such systems should be evaluated by outcome metrics - feature delivery acceleration, defect rate, time to recovery - rather than output metrics such as pull request counts or the proportion of AI-generated code. The analysis traces an architectural evolution from a conventional sandbox-plus-coding-agent configuration, found "pretty indistinguishable" from engineers running Claude Code locally, toward a two-layer system comprising TARS (interaction and orchestration) and Horizon (infrastructure orchestration fronting an MCP gateway). The central finding is that durable leverage arises from encoding the entire product engineering process - specification authoring, ticket decomposition, triage, and support - into automation, not from code generation alone. Practical implications include prioritizing an internal MCP gateway as a first investment and building an evergreen organizational memory layer to support continuous system improvement.
1. Introduction
The term software factory denotes an automated pipeline in which a cloud sandbox, an autonomous coding agent, and a natural-language prompt collectively produce a merge-ready pull request. The pattern gained industry attention following Ramp's publication describing its internal "inspects" system, which, as one practitioner observed, "really excited the entire industry of thinking like, oh, we can actually build software very differently now."
This synthesis interrogates a question the initial enthusiasm largely deferred: what should a software factory be measured by, and what must it automate to justify its existence? The central thesis is that output metrics - the number or percentage of agent-authored PRs, lines of AI-generated code shipped - are weak proxies for value, since elevated PR volume need not correspond to accelerated product delivery. Outcome metrics, by contrast, capture the rate at which customer-facing features reach production and the quality with which they do so.
The analysis proceeds by establishing the intellectual context (Section 2), presenting core architectural and process findings organized around system decomposition, process encoding, context infrastructure, and an illustrative workflow (Section 3), distilling actionable technical insights (Section 4), and discussing broader implications before concluding.
2. Background and Related Work
The canonical software factory configuration consists of three components: an isolated execution sandbox, a coding agent such as Claude Code or Open Code, and a prompt interface that triggers a run terminating in a PR. Ramp's published system served as industry precedent and motivated comparable efforts elsewhere.
WorkOS initially replicated this pattern, implementing sandboxes on Cloudflare infrastructure with an open-code model router. The empirical result was negative relative to the organization's goals: the system delivered "not an incremental increase or an exponential increase over engineers just driving cloud code on their laptops. It was actually pretty indistinguishable." This finding motivates a reframing: if the marginal value of a hosted sandbox over a local agent harness is near zero, the sandbox is not the locus of leverage. Leverage must instead reside in the surrounding engineering process - specification, coordination, triage, and review - that precedes and follows code authorship. A second relevant framework is product engineering culture, in which engineers absorb product management responsibilities directly, a structural condition at WorkOS that shapes what the factory must automate.
3. Core Analysis
3.1 System Decomposition: TARS and Horizon
WorkOS restructured its factory into two cooperating layers. TARS serves as the user interaction layer, subscribing to webhooks from Slack, Linear, and GitHub to track project progress rather than merely generating code on demand. Horizon functions as the infrastructure orchestration layer, positioned in front of an MCP gateway in a manner structurally similar to systems like Ramp's inspects or comparable "minions" architectures. This separation decouples the question of what work should happen next (TARS) from where and how that work executes (Horizon), allowing the orchestration substrate to evolve independently of the interaction surface that engineers and stakeholders actually experience in Slack and Linear.
3.2 Encoding Product Engineering Processes
The second major finding concerns process encoding rather than code generation. WorkOS operates without dedicated product managers, relying instead on hilltop documents - PRDs covering purpose, customer feedback, competitive analysis, design screens, and milestones. A dedicated "PM" agent drafts the first version of each hilltop document, incorporates contextual research, and decomposes the plan into Linear tickets with explicit dependencies. This dependency graph is functionally significant: it allows TARS to autonomously pick up subsequent tasks within a cycle as prior tickets complete, and to reevaluate the Linear project over time to detect gaps as work progresses. Humans remain in the loop at the review and refinement stage, editing AI-generated tickets and documents rather than authoring them from scratch. As stated directly, "it's not sufficient that the factory is producing code. We actually want it to take a lot of the other work that our engineers do to produce products and automate that as well."
3.3 Illustrative Workflow: Vaults API
The Vaults API project demonstrates the end-to-end pipeline. The project was initiated via a brief Slack command containing a few sentences of description. TARS auto-generated a Linear project, draft Notion documents, a decision log, and a set of open questions. An engineer then refined the specification and returned it to TARS for execution. Notably, the agent was observed to overestimate scope on occasion, requiring human intervention to cut back the plan - an explicit limitation rather than a claimed success. The resulting documentation was not restricted to TARS's own execution: other agents, including Devin and local Claude Code/Opus harnesses, were able to pick up the same tickets using identical context retrieved via MCP, with security and other collaborating teams weighing in within the project's communication channel.
3.4 Context Infrastructure: The MCP Gateway
Underlying both the PM agent and TARS's triage capabilities is an internal MCP gateway referred to as the context engine, which connects to internal systems such as Snowflake. Semantic tables within this gateway describe product utilization and customer conversation data, and the MCP server itself supplies tool descriptions and guidance on which tables or tools are appropriate for a given query. Although originally constructed to connect the coding agent to internal systems, the gateway is now used more broadly for internal data and customer analysis conducted directly through Slack. This generalization is treated as the single most important recommendation for teams beginning a software factory initiative: building an internal MCP gateway with well-specified tool descriptions and organizational context before investing further in agent orchestration.
4. Technical Insights
Several implementation-level findings merit explicit attention for practitioners designing comparable systems.
- Webhook-driven state tracking: TARS's integration with Slack, Linear, and GitHub via webhooks allows it to maintain situational awareness of project state without requiring a dedicated polling or prompt-triggered invocation model.
- Dependency-aware ticket execution: Encoding ticket dependencies directly in Linear enables autonomous sequential task pickup, reducing the coordination burden otherwise placed on human engineers managing a cycle.
- Agent interchangeability via shared context: Because documentation and context are exposed through MCP rather than embedded in a single agent's internal state, multiple coding agents (TARS, Devin, local harnesses) can operate against the same project artifacts, reducing lock-in to any single orchestration system.
- Semantic data layers over raw access: The Snowflake-backed context engine does not merely expose raw tables; it provides semantic descriptions and usage guidance, which is presented as a precondition for reliable agent query behavior.
- Infrastructure ownership as a strategic goal: WorkOS is using TARS to build its own sandbox infrastructure, moving off third-party sandbox-as-a-service providers to gain deeper control over session information and the ability to migrate workloads across infrastructure.
- Trade-off acknowledgment: The scope-overestimation behavior observed in the Vaults API workflow indicates that autonomous planning still requires human correction, limiting the degree of unsupervised execution currently achievable.
5. Discussion
These findings suggest a broader reframing of what "AI engineering infrastructure" should optimize for. The negative result from WorkOS's initial sandbox implementation - indistinguishability from local agent usage - is a useful null result rarely reported in industry discourse, which tends to emphasize successful deployments. It implies that sandbox infrastructure alone, regardless of execution environment sophistication, does not constitute the differentiating layer of a software factory; the differentiating layer is process and context encoding.
This has implications for how organizations should sequence investment. Rather than building increasingly sophisticated sandboxing or model-routing infrastructure, the evidence here points toward prioritizing context infrastructure (the MCP gateway) and process encoding (hilltop documents, dependency-aware tickets) as prerequisites. An open question concerns generalizability: WorkOS's product engineering culture, which lacks dedicated PMs, may make the PM-agent role more structurally necessary than it would be at organizations with existing product management functions. Future work could examine whether the hilltop-document approach transfers to organizations with different role structures.
The planned evergreen memory layer, capturing per-person, per-team, and org-level semantic context, represents an unverified future direction rather than a demonstrated result. Its stated purpose - enabling continuous, semi-online identification of skill gaps and outdated practices - suggests an ambition toward self-improving infrastructure, though no outcome data is yet available to substantiate this trajectory.
6. Conclusion
This synthesis contributes an empirically grounded counterpoint to output-metric-driven evaluation of software factories, demonstrating through WorkOS's experience that sandbox-and-agent systems alone fail to outperform local agent usage, and that genuine leverage requires encoding product engineering process - specification, ticketing, triage, and support - into automation via systems like TARS, Horizon, and an MCP-based context engine. The practical takeaway for teams beginning similar initiatives is to invest first in internal context infrastructure rather than execution environments, and to measure success through defect rate, time to recovery, and organic engineer adoption rather than PR counts or generated-code volume.
Sources
- No, That's Not a Software Factory - Ryan Cooke, WorkOS - Original Creator (YouTube)
- Analysis and summary by Sean Weldon using AI-assisted research tools
About the Author
Sean Weldon is an AI engineer and systems architect specializing in autonomous systems, agentic workflows, and applied machine learning. He builds production AI systems that automate complex business operations.