MCP Doesn't Suck. Your Agent Does. - Jan Čurn, Apify

MCP itself is not fundamentally broken; the problem lies in poor agent/client implementations that misuse the protocol, and combining MCP with CLI-style acce...

By Sean Weldon

Abstract

The Model Context Protocol (MCP) has become the dominant standard for connecting large language model agents to external tools, with an estimated 10,000-15,000 deployed servers. A vocal backlash, however, characterizes MCP as broken, citing excessive context consumption of up to one-third of the available window before task execution begins. This synthesis, drawing on analysis presented by Jan Čurn of Apify, argues that this deficiency originates not in the protocol specification but in agent harness implementations that naively load all tool definitions upfront. Three remediation strategies - sub-agent partitioning, progressive tool discovery, and code mode - are examined alongside a comparative analysis of command-line interfaces (CLIs). This motivates MCPC, a universal CLI client for MCP. Preliminary benchmarking via a "connector evals" framework indicates raw MCP consumes substantially more tokens than CLI or MCPC connectors at comparable completion speed, suggesting hybrid approaches merit broader adoption.

1. Introduction

MCP, introduced by Anthropic approximately two years ago, defines a standardized interface through which autonomous agents discover and invoke external capabilities. Its adoption has been rapid and extensive: thousands of community and vendor-operated servers, multiple competing registries, and first-class integration into major coding agents have emerged within a short period. This growth pattern is characteristic of infrastructure protocols that solve a genuine coordination problem - prior to MCP, each agent framework implemented bespoke, incompatible function-calling schemas.

Despite this adoption curve, prominent practitioners - including Anthropic itself, along with public figures such as Gary Tan and Peter Levels - have openly questioned MCP's viability. The recurring complaint is that MCP "eats context," with tool registration alone consuming up to one-third of an agent's context window before any substantive work occurs. This criticism has been influential enough to generate sentiment that the protocol is fundamentally flawed or in decline.

This analysis advances a specific counter-thesis: MCP is not fundamentally broken; the agent harness is. A harness refers to the client-side implementation layer that manages how an agent interacts with the protocol - determining when tools are loaded into context, how results are retained, and how sessions persist. The protocol specification defines what tools exist and how they are invoked; it is silent on when and how much of that information should enter the model's working context. The central claim is that naive harness design, not protocol design, produces the observed pathology.

The analysis proceeds in three stages: first, identifying the root cause of context bloat and evaluating three proposed remediations; second, comparing MCP against CLI-based tool access to isolate complementary strengths; and third, presenting MCPC, a CLI client for MCP, along with preliminary benchmarking evidence from a "connector evals" framework.

2. Background and Related Work

MCP standardizes three primitives - tools (callable functions), resources (addressable data), and prompts (reusable templates) - over a transport layer supporting both local stdio connections and authenticated remote access. This standardization replaced a fragmented landscape of proprietary function-calling schemas, enabling interoperability across agent frameworks and tool providers.

The criticism directed at MCP is best understood as a critique of context economics: every tool definition registered with a model occupies tokens, incurring monetary cost, latency, and attention dilution that can degrade retrieval accuracy. A related but distinct failure mode concerns data hygiene - sensitive values such as passwords, once returned by a tool call, persist in conversation history and remain exposed to subsequent unrelated calls. Three remediation frameworks have emerged in response: sub-agent context delegation, Anthropic's tool search tool (subsequently adopted by Cursor), and Cloudflare's code mode. Each is evaluated below.

3. Core Analysis

3.1 The Harness, Not the Protocol

The MCP specification does not prescribe a context management policy; tool loading, caching, and scoping are implementation concerns left to the client. Harnesses that eagerly register the full tool catalogue of every connected server at session start produce context consumption that scales linearly with the number of available tools, regardless of task relevance - consuming roughly one-third of the context budget before productive work begins. Critically, sub-agent partitioning, often proposed as a fix, reduces pollution of the main context but does not eliminate aggregate token cost, nor does it resolve the sensitive-data persistence problem, since any sub-agent that retrieves a credential still retains it in its own context.

3.2 Progressive Discovery and Code Mode

Two more substantive remediations target the loading mechanism itself. Anthropic's tool search tool enables progressive tool discovery, loading tool definitions into context only when relevant to the current task, yielding significant token savings; this approach has since been adopted by Cursor. Cloudflare's code mode takes a different approach, treating MCP tools as code artifacts to be written and executed rather than as structured function calls loaded into context. This is effective because, as the analysis notes, "tool calling is an artificial construct that we have to teach the LLMs to do - it doesn't exist in a real world in real training data," whereas code is extensively represented in training corpora, giving models stronger native proficiency. The limitation of Cloudflare's implementation is practical rather than conceptual: it is platform-specific, constraining broader adoption.

3.3 CLI as an Implicit Code-Mode Baseline

A comparative analysis of CLI-based tool access reveals why command-line interfaces largely avoid the context bloat problem by default. Agents never load full CLI documentation into context; instead, they invoke commands progressively and already possess extensive command knowledge from training data. Unix/Linux shell commands, dating to 1969 (Ken Thompson and Dennis Ritchie), are heavily represented in training corpora, such that "agents know the shell by heart." Furthermore, CLI tools execute as code by default - they require a sandboxed runtime or machine - meaning CLI access possesses "code mode" characteristics inherently, whereas MCP had to evolve toward this property through separate mechanisms like Cloudflare's implementation.

CLI's limitation is architectural: it functions as a local black box lacking a standardized transport protocol, which complicates remote access, instrumentation, and credential injection - precisely the problems MCP was designed to solve. The resulting conclusion is complementary rather than oppositional: MCP is suited to standardized remote access, while CLI is suited to local agent interfaces, and combining both captures the respective benefits of each.

4. Technical Insights

Several implementation-level findings carry direct engineering relevance:

Implementation trade-offs remain: asynchronous task support, while valuable for long-running operations, is inconsistently implemented across MCP servers, and the evals framework itself is described as early-stage and open to community contribution, limiting the current statistical robustness of its conclusions.

5. Discussion

The broader implication of this analysis is that debates framing MCP as categorically "broken" conflate protocol design with implementation quality. This distinction matters practically: discarding MCP in favor of ad hoc alternatives would forfeit its genuine contributions - standardized remote transport, authentication, and cross-vendor interoperability - without addressing the underlying harness deficiencies that caused the criticism in the first place.

The CLI/MCP comparison further suggests that the industry's pursuit of code mode as a novel innovation overlooks that CLI tooling has implicitly satisfied this property for decades, by virtue of executing within sandboxed runtimes and benefiting from five decades of training-data representation. This raises a broader question for future agent architecture: rather than retrofitting code-mode behavior onto context-loaded protocols, hybrid designs that route local operations through CLI-like execution and remote operations through MCP-like transport may be architecturally preferable.

Open questions remain regarding the generalizability of the connector evals results beyond the Claude Code/Sonnet 5 configuration tested, and regarding how sensitive-data persistence should be formally addressed at the protocol or harness level, since none of the three remediation strategies examined fully resolves it.

6. Conclusion

This analysis contributes a reframing of the MCP backlash: the protocol's specification is not the source of observed inefficiencies; harness implementation choices are. Progressive tool discovery and code mode both offer partial remediation, and the structural similarities between CLI execution and code mode suggest that CLI-style access, combined with MCP's standardized transport, constitutes a pragmatic hybrid. MCPC operationalizes this hybrid as a protocol-agnostic CLI client, and early benchmarking indicates measurable token-efficiency advantages over raw MCP without sacrificing completion speed. Practitioners building agent harnesses should prioritize deferred tool loading and composable, scriptable interfaces over eager, context-exhaustive registration as a default design pattern.


Sources


About the Author

Sean Weldon is an AI engineer and systems architect specializing in autonomous systems, agentic workflows, and applied machine learning. He builds production AI systems that automate complex business operations.

LinkedIn | Website | GitHub