What It Actually Takes to Build a Software Factory - Tereza Tížková, Factory
A software factory is the full autonomous lifecycle of software development - collecting signals, prioritizing, orchestrating, executing, validating, and conti...
By Sean WeldonWhat It Actually Takes to Build a Software Factory
Abstract
The term software factory has become prominent in discussions of agentic software engineering, yet it is often conflated with coding agents or multi-agent systems. This synthesis, based on practitioner analysis from Tereza Tížková of Factory, argues that a software factory constitutes the complete autonomous software development lifecycle - signal collection, prioritization, orchestration, execution, validation, and iteration - of which code generation is the least difficult component. Three design principles are identified as prerequisites: remaining agnostic to models and tools, achieving genuine autonomy through well-defined completion criteria, and continuously improving through disciplined context management. Reported metrics include approximately 25% cost reduction from automatic model routing, 50%+ token savings from deferred context disclosure, and validation processes consuming up to 40% of mission runtime. Implications for codebase readiness and the evolving human role in software development are discussed.
1. Introduction
The rapid expansion of agentic AI capability has generated widespread interest in autonomous software development, yet practical implementation experience remains concentrated among a small set of practitioners. This gap between discourse and execution is captured directly: "Everyone is talking about software factory but only few people are actually building one and even fewer people know what it actually takes to build one."
This synthesis defines the software factory construct and distinguishes it from adjacent, frequently conflated concepts. A software factory is understood here as the full lifecycle of software development executed with autonomy: collecting signals from running systems, reacting to user feedback and application logs, prioritizing work, orchestrating agent teams, executing changes, validating outputs, testing in production, iterating, and converting outcomes into reusable organizational knowledge.
Critically, a software factory is not a coding agent, nor a swarm of coding agents, regardless of scale. As stated plainly: "Generating code, writing code - that's the easy part compared to all the others." Nor is it a consultancy engagement or strategic framework layered atop existing structures; the argument advanced is that organizations must rebuild from the ground up rather than insert agentic tooling into unchanged workflows.
The analysis proceeds through historical feasibility conditions (§2), three core design principles - agnosticism, autonomy, and continuous improvement (§3) - consolidated technical findings (§4), and implications for organizational readiness and human roles (§5).
2. Background and Related Work
The infeasibility of software factories prior to approximately 2023 stemmed from four concrete limitations in contemporaneous language models: high hallucination rates, insufficient context length for sustained reasoning over codebases, inadequate reasoning quality for multi-step planning, and the absence of isolated execution environments in which agents could safely act and observe consequences.
These constraints have since relaxed substantially, particularly through advances in computer use and persistent virtual environments, which allow agents to interactively exercise software rather than produce artifacts that merely appear correct. Notably, iterative agent loop patterns - such as the Ralph loop - are not themselves novel innovations; their existence predates current systems. The genuine contribution lies not in the looping mechanism but in solving the harder problem of defining "done" for open-ended, nondeterministic engineering tasks.
3. Core Analysis
3.1 Being Agnostic
A software factory must remain agnostic to environments (e.g., Slack, GitHub) and to existing model subscriptions, since frontier models change continuously. This principle is illustrated through a Coinbase case in which AI spend was reduced without reducing token spend, achieved through default model selection, caching, and routing rather than usage restriction.
The mechanism underlying this is automatic model routing: a system that classifies task difficulty, sets a performance threshold, and selects the cheapest model capable of meeting that threshold - switching models automatically upon failure. This yields benefits beyond cost, including improved reliability and speed. Benchmarking indicates a conservative 25% cost savings, with likely higher real-world gains. Importantly, caching is not restricted to proprietary model providers; open models running on dedicated compute can achieve comparable caching benefits, meaning final pricing to end users reflects business decisions rather than fixed technical constraints.
3.2 Autonomy and Mission Structure
Genuine autonomy requires solving the problem of defining completion criteria for nondeterministic tasks. Poorly specified criteria allow agents to satisfy tests technically without accomplishing the underlying goal - a failure mode described as agents "cheating" the validation process.
This is addressed through factory missions: long-running agent sessions, observed extending from 16 hours to several weeks, structured around orchestrator, worker, and validator agents. Worker agents operate sequentially rather than as parallel swarms, preserving fresh context in a manner analogous to human code-review handoffs; each sequential worker may nonetheless spawn parallel sub-agents for bounded subtasks.
Validators assess code they did not author, guided by a validation contract authored upfront by the orchestrator. Two validator types are distinguished: a scrutiny validator, which evaluates code quality, linting, and test coverage, and a user testing validator, which conducts interactive testing within a virtual environment. In a documented 16-hour customer mission, validation consumed up to 40% of total process time - underscoring that verification, not generation, is the dominant cost center in autonomous pipelines. This capability is itself contingent on recent progress in computer-use technology and persistent virtual environments, which allow interactive testing rather than superficial output inspection.
3.3 Continuous Improvement Through Context Management
Enterprise environments typically involve hundreds of tools (Figma, Notion, Gmail, Drive, Slack), producing context bloat and increasing the likelihood of tool misselection by agents. The proposed solution is a deferred context engine, which progressively discloses tools and context only as needed. Information is not deleted but withheld until relevant: "Nothing is actually removed, it's just hidden and not reachable until needed." This approach is reported to save 50% or more tokens at scale as the number of available tools grows.
Continuous improvement also depends on codebase quality. AI adoption is characterized as following a power law: well-structured, documented codebases benefit disproportionately, while unstructured codebases degrade further under AI-driven modification. Cited Stanford data supports the claim that AI can worsen code quality absent proper structure and documentation: "If your codebase is not ready... then adopting or turning into software factory can actually make you end up worse and make your code degrading." This motivates an agent readiness framework - a hygiene check spanning reproducible environments, test coverage, documentation, and code style consistency. Plugins, defined as packaged reusable skills and context, codify implicit team norms and support automatically updated documentation as a mechanism for sustaining readiness over time.
4. Technical Insights
Several quantitative findings carry direct implementation relevance. The 25% cost savings from automatic model routing is described as conservative, suggesting routing-based architectures warrant consideration even in cost-insensitive deployments, given secondary gains in reliability and latency. The 50%+ token savings from deferred context disclosure scales with tool count, implying the benefit compounds as enterprise tool ecosystems expand rather than remaining fixed.
The finding that validation consumes up to 40% of mission runtime has architectural implications: systems optimized purely for generation speed may be addressing a minority of total latency. This suggests validator design - particularly the separation of scrutiny and interactive user-testing validators - merits equal engineering investment as execution agents.
The sequential-worker-with-parallel-sub-agent pattern represents a trade-off between context freshness and raw parallelism, prioritizing coherence over throughput for long-running tasks. This contrasts with swarm-based architectures and suggests that mission duration (hours to weeks) rather than task count is the relevant scaling variable for orchestration design. Finally, the compatibility of caching with open models on dedicated compute indicates that cost advantages attributed to specific vendors may be infrastructural rather than proprietary, with pricing differentiation functioning as a business rather than technical decision.
5. Discussion
These findings collectively suggest that software factory construction is less a modeling problem than a systems-engineering and organizational-readiness problem. The disproportionate time allocated to validation, and the power-law relationship between codebase hygiene and AI-driven outcomes, indicate that technical capability alone is insufficient without parallel investment in documentation, test coverage, and environment reproducibility.
A notable tension emerges between autonomy and oversight: achieving genuine autonomy requires solving validation rigorously enough that human review becomes unnecessary, yet the resources required to build robust validator agents are substantial. This suggests near-term software factories may remain validation-heavy by design rather than as a transitional inefficiency.
The historical framing - human computers to programming languages to coding agents to software factories - positions the current moment as a continuation of abstraction-layer progression rather than a discontinuity. The thesis that humans should determine what to build rather than how carries implications for workforce role redefinition, with AI absorbing coordination overhead (status syncs, alignment meetings) rather than creative or architectural work.
6. Conclusion
This synthesis establishes that a software factory is best understood as a complete autonomous development lifecycle rather than a coding agent deployment, with validation, context management, and model-agnostic routing constituting the primary engineering challenges. Reported metrics - 25% routing-based cost savings, 50%+ context-deferral token savings, and 40% validation time allocation - indicate where implementation effort should concentrate.
Practically, organizations considering software factory adoption should first assess agent readiness across documentation, testing, and environment reproducibility, since the evidence suggests AI adoption without such foundations risks degrading rather than improving code quality.
Sources
- What It Actually Takes to Build a Software Factory - Tereza Tížková, Factory - Original Creator (YouTube)
- Analysis and summary by Sean Weldon using AI-assisted research tools
About the Author
Sean Weldon is an AI engineer and systems architect specializing in autonomous systems, agentic workflows, and applied machine learning. He builds production AI systems that automate complex business operations.