'Your Agents Are in Solitary Confinement: Why MCP & A2A Aren''t Enough - Vlad Luzin, Band'

The future belongs to fully autonomous AI-to-AI communication across businesses and consumers, but achieving this requires solving hard distributed systems p...

By Sean Weldon

Your Agents Are in Solitary Confinement: Why MCP & A2A Aren't Enough

Abstract

This synthesis examines the technical argument that autonomous agent-to-agent (A2A) communication - within enterprises, between enterprises, and between consumers and businesses - constitutes the next substantive layer of AI infrastructure, and that existing protocols fail to deliver it. The analysis characterizes multi-agent deployments as distributed systems of non-deterministic microservices, identifying four unresolved infrastructure requirements: an ordered, retry-capable transport layer; continuity and state hydration; runtime binding of thread, conversation, and execution identifiers; and a governance layer supplying identity and audit. The Model Context Protocol (MCP) is shown to be stateless and therefore unsuitable for sticky sessions, while the A2A protocol is client-server and omits discovery. Two systems, Band (a collaboration layer) and Jam (a desktop orchestration client), are presented as implementations that raise abstraction to conversational primitives - rooms, channels, participants - with bilateral consent, cost attribution, and real-time observability.

1. Introduction

The prevailing deployment pattern for large language model (LLM) agents remains fundamentally human-mediated. A developer may operate several concurrent coding sessions - multiple Claude or Codex instances - while personally copying output from one session into the prompt of another. In this configuration the human functions as a message router between stateful processes that have no channel to address one another directly. The condition has been characterized bluntly:

"Your agent is still alone. They cannot see each other. They cannot communicate with each other. They are in digital solitary confinement."

The central thesis advanced here is that the future of applied AI belongs to fully autonomous AI-to-AI communication, and that realizing this future is not a prompting problem but a distributed systems problem. Agents will be authored in heterogeneous frameworks and languages, deployed across divergent runtime environments, and expected to discover one another, form conversational spaces, negotiate tasks, and report results to humans without intervening manual routing. Current protocols address fragments of this requirement; none address the whole.

Key terminology for this analysis includes loop engineering, the practice of configuring agents to prompt one another automatically rather than relying on human routing, and the four infrastructure layers identified as necessary but absent from current tooling: transport, continuity, runtime binding, and governance. This paper proceeds by establishing the intellectual context of agent interoperability protocols (Section 2), analyzing the specific technical deficits of existing approaches and the infrastructure primitives required to close them (Section 3), extracting actionable engineering insights (Section 4), and discussing broader implications and open questions (Sections 5-6).

2. Background and Related Work

Three protocol efforts frame the current interoperability landscape. MCP (Model Context Protocol) standardizes tool invocation, allowing a model to call external capabilities through a uniform interface. A2A (Agent-to-Agent) protocol specifies interaction between agents as distinct entities rather than as tools. ACP represents an additional entrant in the agent communication protocol space.

Alongside protocol work, loop engineering has emerged as a practical workaround: rather than a human manually routing messages, agents are configured to prompt one another in automated loops. Loop engineering correctly identifies the inefficiency of human-in-the-loop routing, but the analysis contends that it relocates rather than resolves the difficulty - implementers end up contending with Python and TypeScript abstraction libraries instead of addressing the underlying coordination problem. A stated methodological preference frames the intended alternative:

"Before thinking about technical solutions to hypothetical problems multi-agent coordination, I try to keep things simple and avoid the problem."

The implication is that infrastructure should eliminate classes of coordination problems rather than supply additional abstraction layers on top of them.

3. Core Analysis

3.1 Protocol-Level Deficiencies

MCP is architecturally stateless, which makes it well-suited to discrete tool calls but unsuitable for sticky sessions between agents that must maintain shared context across multiple turns. A2A, by contrast, is a client-server protocol; achieving bidirectional communication between two agents requires each to implement both client and server roles simultaneously, effectively doubling the implementation surface. Discovery - the mechanism by which an agent locates and identifies a counterpart - is not part of the A2A specification at all, leaving a foundational capability for autonomous interaction unaddressed. Chaining multiple agents through REST APIs further introduces timeout failures under extended multi-step workflows, a problem that only persistent, ordered queues with delivery tracking can resolve.

3.2 The Distributed Systems Problem

The analysis frames a multi-agent deployment where every agent is remote as, functionally, a distributed system composed of microservices - except that each "microservice" is a non-deterministic model rather than deterministic code:

"A multi-agent system where every agent is remote is a distributed system of microservices where each microservice is non-deterministic."

This reframing is significant because it imports decades of distributed systems concerns - ordering, retries, idempotency, consistency - into a domain where practitioners have often treated coordination as a prompting exercise. Four infrastructure layers are identified as prerequisite: (1) a transport layer guaranteeing ordered, real-time delivery with retries; (2) continuity and consistency through state hydration; (3) runtime binding, which maps thread IDs, conversation IDs, and execution IDs across heterogeneous agent runtimes; and (4) a governance layer providing identity and audit, required even after transport and binding are solved.

3.3 Onboarding Friction as Evidence of the Gap

Empirical evidence for the absence of a unifying layer is drawn from the manual steps required to connect agents to existing messaging platforms: five steps for Telegram, seven for Discord, eight for Slack, and eleven for WhatsApp. This escalating step count across platforms illustrates that, in the absence of a shared abstraction, every integration is bespoke, and the effort compounds with each additional surface an organization wishes to support. The proposed resolution is to raise the abstraction of the technical stack to conversational primitives:

"We need to raise abstraction of a technical stack to conversation and talk about rooms, channels, participants."

3.4 Band and Jam as Implementations

Band is positioned as a global collaboration layer intended to connect any agent, regardless of framework or deployment environment, by implementing the primitives identified above: registration, onboarding, bilateral consent for contact requests, and invitation into shared conversations. A demonstration showed programmatic registration and onboarding of Codex and "Land Rush" agents, followed by cross-user, cross-registry communication gated by mutual consent before any interaction occurred - directly addressing the governance requirement.

Jam, a desktop application built atop Band, targets the onboarding of local agents and addresses routing, context overload, cost attribution, and mixed multi-agent/multi-human collaboration. It captures tasks generated by Claude and Codex in real time, visualizes the software architecture and which components are under active work, and allows remote humans or specialized agents (for example, a security-focused agent) to join a session without duplicating existing skills.

4. Technical Insights

Several implementation-relevant findings emerge from this analysis:

A trade-off worth noting is that raising abstraction to conversational primitives (rooms, channels, participants) simplifies developer experience but requires the underlying platform to absorb the full complexity of transport, binding, and governance - complexity that does not disappear but is relocated into infrastructure.

5. Discussion

The broader implication of this analysis is that multi-agent coordination is being treated, in much of current practice, as a software engineering problem to be solved with orchestration libraries, when the evidence presented suggests it is better understood as an infrastructure problem analogous to telecommunications or messaging systems. The repeated analogy to rooms, channels, and participants signals a deliberate borrowing from messaging platform design rather than from traditional RPC or microservice patterns, reflecting the judgment that agents behave more like conversational participants than like deterministic services.

A notable gap concerns standardization: with MCP, A2A, and ACP each addressing partial slices of the problem, and platforms like Band proposing a unifying layer, the field lacks consensus on whether convergence will occur around an open protocol or around proprietary collaboration layers. The onboarding-step disparity across Telegram, Discord, Slack, and WhatsApp (5, 7, 8, and 11 steps respectively) is suggestive but anecdotal, and further quantitative study of integration effort across a wider set of platforms and protocols would strengthen the case for a unifying abstraction.

This direction connects to a broader industry trend toward treating agents as first-class networked entities with identity, consent, and audit requirements, rather than as embedded functions within a single application. The emphasis on bilateral consent before cross-registry interaction foreshadows governance questions - analogous to federated identity and access management - that will likely intensify as agent populations scale beyond single organizations.

6. Conclusion

This analysis has argued that existing agent interoperability protocols, MCP and A2A chief among them, solve adjacent but insufficient problems: tool invocation and point-to-point client-server messaging, respectively, while leaving transport reliability, state continuity, runtime identifier binding, discovery, and governance unresolved. The practical consequence, evidenced by manual onboarding overhead and timeout-prone chaining, is that multi-agent systems today function as fragile, hand-assembled distributed systems rather than as a coherent platform.

The proposed path forward - exemplified by Band and Jam - is to raise the level of abstraction to conversational primitives and to implement consent, identity, and cost-attribution mechanisms as native platform features rather than bespoke integrations. For practitioners building multi-agent systems, the practical takeaway is to evaluate whether a given protocol choice addresses the full stack of transport, continuity, binding, and governance, or merely one layer of it, before committing to an architecture.


Sources


About the Author

Sean Weldon is an AI engineer and systems architect specializing in autonomous systems, agentic workflows, and applied machine learning. He builds production AI systems that automate complex business operations.

LinkedIn | Website | GitHub