Stop Rationing Tokens: Let the Harness Pick the Model - Kimchi by Cast AI
Instead of rationing tokens to control LLM coding costs, enterprises should use an automated harness that selects the optimal model per task based on outcome...
By Sean WeldonStop Rationing Tokens: Let the Harness Pick the Model
Abstract
Enterprise adoption of large language model (LLM) coding agents has produced cost curves that outpace budget planning, prompting widespread token rationing. This synthesis examines an alternative: treating token volume as effectively unlimited while controlling expenditure through automated, outcome-aware model selection. Drawing on internal deployment data from Kimchi, a 300-person engineering organization, the analysis contrasts cost-per-token with cost-per-task as the operative economic unit, showing that nominally inexpensive models can be substantially more expensive once task completion is measured. A multi-model harness that scores artifact quality and routes work accordingly yielded a reported 2.5x reduction against projected cloud spend, with token usage rising 1.5x while unit coding cost fell 1.5x. Supporting infrastructure for remote session execution (Teleport) and collaborative review (Studio) is examined, alongside a broader shift from diff-level code review toward specification-level review.
1. Introduction
Agentic coding tools have converted software development into a high-throughput consumer of inference capacity, and enterprise budgeting has not kept pace. The scale of this mismatch is illustrated by two reported figures: one Indian company is said to have spent $500 million on Anthropic services in a single month, while Uber's CTO is reported to have exhausted an entire year's Anthropic budget in four months. The typical organizational response has been to impose per-developer token quotas - a policy this analysis characterizes as structurally flawed, analogous to permitting a laptop to be charged only once per day.
The central terminological distinction underpinning this synthesis is between cost-per-token, the conventional unit used in model pricing and procurement, and cost-per-task, a unit that normalizes expenditure against completed work rather than raw consumption. The thesis advanced is that an automated harness capable of selecting the optimal model per task, based on measured outcome quality, can decouple cost growth from token growth. Under this framing, tokens become effectively unlimited while the cost of producing working software decreases.
This analysis proceeds by examining the empirical basis for the cost-per-task distinction, the architecture and measured performance of Kimchi's multi-model harness (Ferment), and the supporting infrastructure - Teleport and Studio - that enables this model of unconstrained, review-governed agentic development. Technical insights and broader industry implications follow.
2. Background and Related Work
The argument rests on two inputs. The first is a university white paper establishing that the true cost of LLM usage diverges materially depending on whether it is normalized per token or per completed task. This finding undermines the prevailing procurement heuristic of selecting models by published per-million-token rates, since token price is shown to be a poor proxy for the price of delivered work.
The second input is the deployment context itself. Kimchi, with approximately 300 employees (roughly two-thirds developers), operated its own coding agent internally for three months prior to the measurements reported here. This dogfooding period provides the longitudinal data underlying the cost and model-selection findings. The organizing philosophy is stated directly: "Our job is not to prevent the developer to use coding agent. Our job is to make sure they can use it as much as they want for as long as they want in a completely unlimited fashion." This reframes managerial responsibility away from gatekeeping consumption and toward engineering a substrate that makes consumption cheap.
3. Core Analysis
3.1 The Divergence of Token Price and Task Price
The empirical centerpiece is the demonstration that models priced similarly per token can diverge sharply per task. Gemini 3 Flash, priced at $3.5 per million tokens on a blended average, costs $75 per task when measured end-to-end. Minimax 2.7, priced at $1.5 per token - nominally cheaper - costs $148 per task, roughly double Gemini 3 Flash's per-task cost despite its lower token rate. As the source material states: "It means the models are not the same. They don't cost the same, but they also not the same per task." This divergence arises because cheaper models may require more retries, longer reasoning chains, or additional correction cycles to reach an acceptable output, inflating the effective cost of a completed unit of work even when per-token pricing is favorable.
3.2 Multi-Model Harness: Measured Outcomes
To exploit this divergence, Kimchi built an automated harness that selects a model per task based on outcome quality rather than static pricing. Over the measured period, this produced a 2.5x savings relative to the projected cloud bill. Notably, token usage itself increased by 1.5x over the same period, while the cost to code decreased by 1.5x - implying a roughly threefold improvement in cost-efficiency relative to naive extrapolation from token growth. This result is significant because it decouples two variables - token volume and total spend - that are conventionally assumed to move in lockstep.
A further finding concerns the harness's responsiveness to the model release cycle. Model selection shifted abruptly on June 12th, moving from Kim 2.6 dominance to Minimax 3 dominance. This shift reflects the harness's capacity to re-evaluate and re-route automatically upon new model availability, a capability framed explicitly against human benchmarking cycles: "Do you think we look at this as human? Absolutely not. But an autonomous harness, an automated coding agent that is obsessed with cost is going to do just that." The implication is that human-driven model evaluation introduces latency that an automated, cost-obsessed routing layer eliminates.
3.3 Ferment: Harness Architecture
The harness, named Ferment, is built atop an open-source SDK (referenced as Pono SDK) and is designed for long-running tasks requiring human check-ins only every two to three hours. Work is decomposed into milestones executed autonomously, with a scoring mechanism validating output quality before model-selection decisions are considered successful. Artifacts must achieve at least a B grade before being treated as complete, with an option to request further refinement toward an A grade. The harness encompasses a full software development lifecycle (SDLC): build, test, a fix-and-rebuild loop, and deployment to staging. Production deployment currently retains a human-in-the-loop gate, though full automation of this final step is identified as a future objective.
3.4 Supporting Infrastructure: Teleport and Studio
Two auxiliary systems extend the harness's practical viability. Kimchi Teleport allows coding sessions to execute in secure remote sandboxes (e.g., on Google Cloud), decoupled from the developer's local machine. Its origin is traced to an incident in which a developer's laptop battery died mid-session; the environment now synchronizes with a container running in a Kubernetes cluster on a hyperscaler, allowing sessions to persist through laptop closure, travel, or leave. Reportedly, 62% of engineers now use Teleport exclusively for coding tasks.
Kimchi Studio extends this capability to team and enterprise collaboration, providing a dashboard visualizing running sessions, including those requiring review, organized in a Kanban-style workflow (backlog, in progress, in review). This structure allows project managers, peer engineers, or non-technical stakeholders to review plans and specifications prior to execution, addressing the difficulty of reviewing large, low-legibility pull requests (e.g., exceeding 2,000 lines) by shifting review emphasis from diff inspection to specification and intent.
4. Technical Insights
Several implementation-relevant findings emerge. First, model selection should be driven by a scoring function applied to task artifacts rather than static per-token pricing tables; the B-grade minimum threshold in Ferment illustrates a practical mechanism for enforcing output quality while permitting cost-optimized model routing. Second, harness architectures benefit from decomposing long-running agentic work into milestones with bounded human-review intervals (here, 2-3 hours), balancing autonomy against oversight. Third, remote execution infrastructure (Teleport) is a prerequisite for sustained autonomous operation, since local-machine dependency introduces failure modes unrelated to model performance. A trade-off remains at the production-deployment boundary, where human gating persists as a safety mechanism despite the stated goal of full automation. Organizations adopting similar harnesses should anticipate the need for continuous recalibration logic, since the June 12th shift from Kim 2.6 to Minimax 3 demonstrates that optimal model assignments are not static and must be re-evaluated automatically as new models release.
5. Discussion
These findings suggest that enterprise LLM cost governance is better modeled as an optimization problem over outcome quality than as a constraint problem over resource consumption. The observation that token usage rose 1.5x while cost fell 1.5x indicates that volume and expenditure can be deliberately decoupled through routing intelligence rather than access restriction. This has implications for procurement practices that currently rely on published token pricing as a comparative benchmark; such benchmarks appear insufficient absent task-level cost normalization.
A further implication concerns code review practices. As harness-generated pull requests grow large and diff-level inspection becomes impractical, review processes must shift toward specification and intent, consistent with the note that "Reading code is not enough anymore." Open questions remain regarding the generalizability of the 2.5x savings figure outside Kimchi's specific task mix, and whether scoring-based quality gates (the B-grade threshold) are robust across domains beyond software engineering.
6. Conclusion
This synthesis presents evidence that automated, outcome-aware model selection can reduce coding costs even as token consumption increases, challenging the rationing paradigm common in enterprise LLM governance. The Ferment harness, combined with Teleport and Studio, demonstrates a coherent architecture for unconstrained agentic development governed by quality scoring rather than access limits. Practically, organizations should prioritize task-level cost measurement over token-level pricing comparisons and invest in infrastructure enabling continuous, automated model re-evaluation as the model landscape evolves.
Sources
- Stop Rationing Tokens: Let the Harness Pick the Model - Kimchi by Cast AI - Original Creator (YouTube)
- Analysis and summary by Sean Weldon using AI-assisted research tools
About the Author
Sean Weldon is an AI engineer and systems architect specializing in autonomous systems, agentic workflows, and applied machine learning. He builds production AI systems that automate complex business operations.