{
  "title": "When Agents Talk Too Much: The Hidden Crisis of Multi-Agent Communication Efficiency",
  "url": "https://miaok.ong/en/posts/agent-multi-agent-communication-crisis/",
  "date": "2026-06-07T09:00:00+08:00",
  "lastmod": "2026-06-07T09:00:00+08:00",
  "type": "posts",
  "kind": "page",
  "language": "en",
  "description": "O(n²) communication disaster in multi-agent systems: arXiv 2606.05304 reveals structured protocols deliver 3-5x efficiency gains, and why the thinking-trace era turns this from optimization to survival.",
  "keywords": null,
  "tags": ["AI","Agent","Multi-Agent","Communication","arXiv"],
  "categories": [],
  "author": "Mark (Miao) Kong",
  "image": "https://miaok.ong/images/avatar.jpg",
  "content": "\u003ch1 id=\"when-agents-talk-too-much-the-hidden-crisis-of-multi-agent-communication-efficiency\"\u003e\n  When Agents Talk Too Much: The Hidden Crisis of Multi-Agent Communication Efficiency\n  \u003ca class=\"heading-link\" href=\"#when-agents-talk-too-much-the-hidden-crisis-of-multi-agent-communication-efficiency\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"Link to heading\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003eLink to heading\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h1\u003e\n\u003cblockquote\u003e\n\u003c/blockquote\u003e\n\u003cp\u003e2026-06-07 | Deep Dive | Based on arXiv 2606.05304 + ICLR 2026 SupervisorAgent + Industry Framework Analysis\u003c/p\u003e\n\u003chr\u003e\n\u003cp\u003eIt\u0026rsquo;s June 2026. Three years have passed since AutoGPT and BabyAGI ignited the multi-agent craze.\u003c/p\u003e\n\u003cp\u003eIn those three years, multi-agent collaboration has evolved from a \u0026ldquo;chain two LLM calls with LangChain\u0026rdquo; experiment into an industrial-grade paradigm where 1,000 agents execute concurrently inside Claude Code. Anthropic\u0026rsquo;s just-released Dynamic Workflows lets users schedule thousands of intelligent agents in a single task. OpenAI Codex breaks down software development into chains of sub-agents. CrewAI, AutoGen, and LangGraph each boast tens of thousands of GitHub stars—we seem to have accepted as default that \u0026ldquo;if one agent can\u0026rsquo;t solve it, throw ten agents at it.\u0026rdquo;\u003c/p\u003e\n\u003cp\u003eBut on June 6, 2026, an arXiv paper from the Singapore University of Technology and Design (SUTD) raised a question that forces the entire direction to be re-examined:\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e\u0026ldquo;What Should Agents Say?\u0026quot;—What exactly should agents say to each other?\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThis deceptively simple question punctures the deepest hidden cost inside multi-agent systems (MAS). And when we combine it with a related paper from ICLR 2026, a communication-pattern analysis of mainstream frameworks, and new variables introduced by the reasoning-model era, we reach a sobering conclusion: \u003cstrong\u003einter-agent \u0026ldquo;chatter\u0026rdquo; is shifting from an optimization target to a survival condition.\u003c/strong\u003e\u003c/p\u003e\n\u003chr\u003e\n\u003ch2 id=\"i-the-physics-of-the-problem-the-on-communication-disaster\"\u003e\n  I. The Physics of the Problem: The O(n²) Communication Disaster\n  \u003ca class=\"heading-link\" href=\"#i-the-physics-of-the-problem-the-on-communication-disaster\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"Link to heading\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003eLink to heading\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h2\u003e\n\u003cp\u003eImagine a sequential pipeline with 5 agents, the most typical agent orchestration pattern in 2026:\u003c/p\u003e\n\u003cpre tabindex=\"0\"\u003e\u003ccode\u003eUser request → Agent A analyzes → Agent B codes → Agent C tests → Agent D deploys → Agent E reports\n\u003c/code\u003e\u003c/pre\u003e\u003cp\u003eUnder \u0026ldquo;unconstrained natural-language communication,\u0026rdquo; here\u0026rsquo;s what each interaction actually looks like:\u003c/p\u003e\n\u003col\u003e\n\u003cli\u003eAgent A receives the user request and calls an LLM. Its output includes: requirement understanding, analysis process, initial solution proposal. Roughly 2,000 tokens.\u003c/li\u003e\n\u003cli\u003eAgent B receives as input: user request + Agent A\u0026rsquo;s full 2,000-token output. It then produces roughly 3,000 tokens (code + explanation).\u003c/li\u003e\n\u003cli\u003eAgent C receives: user request + A\u0026rsquo;s 2,000 + B\u0026rsquo;s 3,000. It produces roughly 2,500 tokens (test report).\u003c/li\u003e\n\u003cli\u003eAgent D receives: user request + A(2,000) + B(3,000) + C(2,500). Plus its own 1,500.\u003c/li\u003e\n\u003cli\u003eAgent E receives: user request + A(2,000) + B(3,000) + C(2,500) + D(1,500). Plus its own 1,000.\u003c/li\u003e\n\u003c/ol\u003e\n\u003cp\u003eCumulative token consumption: each downstream agent processes a \u003cstrong\u003elinearly growing\u003c/strong\u003e context, but the API cost per call = input token price × accumulated context size. When a task involves 20 interaction rounds, the 20th round\u0026rsquo;s input may already span tens of thousands of tokens—and at least 70% of those tokens are redundant.\u003c/p\u003e\n\u003cp\u003eAnd this doesn\u0026rsquo;t even factor in reasoning models. If you swap Agent C for Gemini 3.1 Pro Thinking, its internal chain-of-thought alone could be 8,000 tokens of deliberation—all of which gets shoveled into Agents D and E\u0026rsquo;s shared history.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eThis isn\u0026rsquo;t linear growth. This is compound interest.\u003c/strong\u003e\u003c/p\u003e\n\u003chr\u003e\n\u003ch2 id=\"ii-what-the-pact-paper-found-agents-dont-need-to-write-essays\"\u003e\n  II. What the PACT Paper Found: Agents Don\u0026rsquo;t Need to Write Essays\n  \u003ca class=\"heading-link\" href=\"#ii-what-the-pact-paper-found-agents-dont-need-to-write-essays\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"Link to heading\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003eLink to heading\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h2\u003e\n\u003cp\u003eSUTD\u0026rsquo;s \u003cem\u003eWhat Should Agents Say?\u003c/em\u003e is the first paper to systematically treat \u0026ldquo;inter-agent communication content\u0026rdquo; as an independent variable.\u003c/p\u003e\n\u003ch3 id=\"21-five-communication-strategies-none-dominant\"\u003e\n  2.1 Five Communication Strategies, None Dominant\n  \u003ca class=\"heading-link\" href=\"#21-five-communication-strategies-none-dominant\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"Link to heading\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003eLink to heading\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h3\u003e\n\u003cp\u003eThe research team tested the five most common inter-agent communication patterns across two MAS topologies (Split-evidence and Sequential Pipeline):\u003c/p\u003e\n\u003ctable\u003e\n  \u003cthead\u003e\n      \u003ctr\u003e\n          \u003cth\u003eStrategy\u003c/th\u003e\n          \u003cth\u003eApproach\u003c/th\u003e\n          \u003cth\u003eStrength\u003c/th\u003e\n          \u003cth\u003eFatal Flaw\u003c/th\u003e\n      \u003c/tr\u003e\n  \u003c/thead\u003e\n  \u003ctbody\u003e\n      \u003ctr\u003e\n          \u003ctd\u003e\u003cstrong\u003eFull Pass-through\u003c/strong\u003e\u003c/td\u003e\n          \u003ctd\u003eUpstream agent\u0026rsquo;s complete output forwarded verbatim\u003c/td\u003e\n          \u003ctd\u003eNo information loss\u003c/td\u003e\n          \u003ctd\u003eExplosive token accumulation\u003c/td\u003e\n      \u003c/tr\u003e\n      \u003ctr\u003e\n          \u003ctd\u003e\u003cstrong\u003eLLM Summary\u003c/strong\u003e\u003c/td\u003e\n          \u003ctd\u003eAnother LLM summarizes before passing\u003c/td\u003e\n          \u003ctd\u003eReduces tokens\u003c/td\u003e\n          \u003ctd\u003eWho does the summary? More cost. And summary quality is unstable\u003c/td\u003e\n      \u003c/tr\u003e\n      \u003ctr\u003e\n          \u003ctd\u003e\u003cstrong\u003eStructured Extraction\u003c/strong\u003e\u003c/td\u003e\n          \u003ctd\u003eExtract only structured fields\u003c/td\u003e\n          \u003ctd\u003eCompact\u003c/td\u003e\n          \u003ctd\u003eRisks omitting contextual details\u003c/td\u003e\n      \u003c/tr\u003e\n      \u003ctr\u003e\n          \u003ctd\u003e\u003cstrong\u003eDebate\u003c/strong\u003e\u003c/td\u003e\n          \u003ctd\u003eAgents rebut each other over multiple rounds until convergence\u003c/td\u003e\n          \u003ctd\u003eImproves answer quality\u003c/td\u003e\n          \u003ctd\u003eHighest token consumption—every rebuttal round accumulates history\u003c/td\u003e\n      \u003c/tr\u003e\n      \u003ctr\u003e\n          \u003ctd\u003e\u003cstrong\u003eVoting\u003c/strong\u003e\u003c/td\u003e\n          \u003ctd\u003eMultiple agents output independently; majority rule\u003c/td\u003e\n          \u003ctd\u003eSimple and robust\u003c/td\u003e\n          \u003ctd\u003eParallel invocation cost is high; doesn\u0026rsquo;t address why parallelism is needed\u003c/td\u003e\n      \u003c/tr\u003e\n  \u003c/tbody\u003e\n\u003c/table\u003e\n\u003cp\u003e\u003cstrong\u003eCore finding: no single fixed strategy is optimal across all scenarios. But one common pattern emerged—in every effective case, what got transmitted was \u0026ldquo;action-state\u0026rdquo; information, not \u0026ldquo;complete natural-language narratives.\u0026rdquo;\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eIn other words: agents need to fill out forms, not write essays.\u003c/p\u003e\n\u003ch3 id=\"22-pact-turning-agent-communication-into-state-updates\"\u003e\n  2.2 PACT: Turning Agent Communication into State Updates\n  \u003ca class=\"heading-link\" href=\"#22-pact-turning-agent-communication-into-state-updates\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"Link to heading\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003eLink to heading\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h3\u003e\n\u003cp\u003eBased on this finding, the research team proposed \u003cstrong\u003ePACT (Protocolized Action-state Communication and Transmission)\u003c/strong\u003e.\u003c/p\u003e\n\u003cp\u003ePACT\u0026rsquo;s core idea is remarkably simple: \u003cstrong\u003etreat inter-agent communication as a \u0026ldquo;public state update\u0026rdquo; problem.\u003c/strong\u003e When each non-terminal agent completes its task, its output isn\u0026rsquo;t pushed directly into shared history. Instead, it first passes through a \u0026ldquo;projection layer\u0026rdquo;—compressed into a compact action-state record. This record contains only three things:\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003e\u003cstrong\u003eWhat was done\u003c/strong\u003e (action)\u003c/li\u003e\n\u003cli\u003e\u003cstrong\u003eWhat state changed\u003c/strong\u003e (state delta)\u003c/li\u003e\n\u003cli\u003e\u003cstrong\u003eWhat downstream needs to know\u003c/strong\u003e (dependency/flag)\u003c/li\u003e\n\u003c/ul\u003e\n\u003cp\u003eOnly then is the compressed record written into shared history for downstream agents to consume.\u003c/p\u003e\n\u003cp\u003eWhat does this sound like? \u003cstrong\u003eGit commit messages.\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe Git community, through decades of evolution, reached consensus: a commit message should be a one-line subject plus a body that describes \u0026ldquo;what was done\u0026rdquo; and \u0026ldquo;why,\u0026rdquo; not the entire diff and the developer\u0026rsquo;s thought process. Nobody wants to read a 3,000-line commit message. Yet multi-agent systems today are doing the equivalent of stuffing every agent\u0026rsquo;s internal diff, thinking notes, and even failed retry logs into the downstream context.\u003c/p\u003e\n\u003ch3 id=\"23-quantified-impact-not-incrementaltransformational\"\u003e\n  2.3 Quantified Impact: Not Incremental—Transformational\n  \u003ca class=\"heading-link\" href=\"#23-quantified-impact-not-incrementaltransformational\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"Link to heading\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003eLink to heading\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h3\u003e\n\u003cp\u003eThe paper\u0026rsquo;s measurements on production-grade code-agent frameworks are compelling:\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003e\u003cstrong\u003eOpenHands (production dev framework)\u003c/strong\u003e: with PACT, the token-per-resolved metric \u003cstrong\u003edropped 10%\u003c/strong\u003e while the task \u003cstrong\u003eresolve rate increased\u003c/strong\u003e. Cheaper \u003cem\u003eand\u003c/em\u003e more accurate.\u003c/li\u003e\n\u003cli\u003e\u003cstrong\u003eSWE-agent\u003c/strong\u003e: while maintaining the same resolve rate, \u003cstrong\u003einput token consumption was halved\u003c/strong\u003e.\u003c/li\u003e\n\u003c/ul\u003e\n\u003cp\u003e\u0026ldquo;Tokens halved, success rate unchanged\u0026rdquo;—in any engineering metric, that\u0026rsquo;s a first-priority outcome.\u003c/p\u003e\n\u003chr\u003e\n\u003ch2 id=\"iii-supervisoragent-a-different-route-same-destination\"\u003e\n  III. SupervisorAgent: A Different Route, Same Destination\n  \u003ca class=\"heading-link\" href=\"#iii-supervisoragent-a-different-route-same-destination\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"Link to heading\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003eLink to heading\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h2\u003e\n\u003cp\u003eAnother paper accepted at ICLR 2026, \u003cem\u003eStop Wasting Your Tokens: Towards Efficient Runtime Multi-Agent Systems\u003c/em\u003e, attacks the same problem from a different angle.\u003c/p\u003e\n\u003cp\u003eSupervisorAgent\u0026rsquo;s approach: \u003cstrong\u003einsert a lightweight \u0026ldquo;supervisor agent\u0026rdquo; at critical interaction nodes.\u003c/strong\u003e Its job isn\u0026rsquo;t participating in business logic—it does three things purely:\u003c/p\u003e\n\u003col\u003e\n\u003cli\u003e\u003cstrong\u003eFilter noise\u003c/strong\u003e: strip out fluff, repetition, and failed-retry logs from upstream agent output\u003c/li\u003e\n\u003cli\u003e\u003cstrong\u003eCorrect errors\u003c/strong\u003e: intercept erroneous information before it contaminates downstream context\u003c/li\u003e\n\u003cli\u003e\u003cstrong\u003ePurify observations\u003c/strong\u003e: compress what\u0026rsquo;s passed into minimal effective information\u003c/li\u003e\n\u003c/ol\u003e\n\u003cp\u003eCrucially, this supervisor agent is triggered by an \u003cstrong\u003eLLM-free context filter\u003c/strong\u003e—meaning it consumes almost zero tokens itself.\u003c/p\u003e\n\u003cp\u003eOn the GAIA benchmark (one of the most challenging evaluations for multi-agent systems), SupervisorAgent reduced \u003cstrong\u003eSmolagent\u0026rsquo;s average token consumption by 29.68%, with zero loss in success rate.\u003c/strong\u003e The same result held across five additional benchmarks—math reasoning, code generation, question answering—and multiple SoTA base models.\u003c/p\u003e\n\u003cp\u003eIf you put PACT and SupervisorAgent side by side, they\u0026rsquo;re saying the same thing: \u003cstrong\u003ethe default mode of agent communication (free-form natural language + full pass-through) is wrong.\u003c/strong\u003e The difference: PACT solves it at the \u003cstrong\u003eprotocol layer\u003c/strong\u003e (defining what agents should transmit), while SupervisorAgent solves it at the \u003cstrong\u003eruntime layer\u003c/strong\u003e (filtering what gets transmitted). The two are fundamentally complementary.\u003c/p\u003e\n\u003chr\u003e\n\u003ch2 id=\"iv-how-current-mainstream-frameworks-handle-communication-a-reality-check\"\u003e\n  IV. How Current Mainstream Frameworks Handle Communication: A Reality Check\n  \u003ca class=\"heading-link\" href=\"#iv-how-current-mainstream-frameworks-handle-communication-a-reality-check\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"Link to heading\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003eLink to heading\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h2\u003e\n\u003cp\u003eThis calls for a hard look at how today\u0026rsquo;s most popular agent frameworks actually handle communication.\u003c/p\u003e\n\u003ch3 id=\"crewai\"\u003e\n  CrewAI\n  \u003ca class=\"heading-link\" href=\"#crewai\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"Link to heading\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003eLink to heading\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h3\u003e\n\u003cp\u003eForces sequential agent execution. What does that mean? The previous agent\u0026rsquo;s complete output, unchanged, becomes the next agent\u0026rsquo;s entire input context. \u003cstrong\u003eThis is an O(n²) token-exponential trap.\u003c/strong\u003e In a 5-agent pipeline, the 5th agent\u0026rsquo;s input context may be 5× larger than the 1st\u0026rsquo;s.\u003c/p\u003e\n\u003ch3 id=\"autogen-microsoft\"\u003e\n  AutoGen (Microsoft)\n  \u003ca class=\"heading-link\" href=\"#autogen-microsoft\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"Link to heading\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003eLink to heading\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h3\u003e\n\u003cp\u003eBased on a group-chat model, defaulting to full conversation history. As turns accumulate, context per API call undergoes \u003cstrong\u003equadratic explosion\u003c/strong\u003e (every agent\u0026rsquo;s every historical utterance is visible to all). Microsoft has recognized the problem and introduced \u003ccode\u003emax_turns\u003c/code\u003e limits and conversation-summarization mechanisms to truncate long tails—but this is essentially \u0026ldquo;replace raw output with summaries,\u0026rdquo; similar to the approaches tested in the paper, and its effectiveness depends on summary quality.\u003c/p\u003e\n\u003ch3 id=\"langgraph\"\u003e\n  LangGraph\n  \u003ca class=\"heading-link\" href=\"#langgraph\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"Link to heading\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003eLink to heading\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h3\u003e\n\u003cp\u003eBased on graphs and state machines. All agents read from and write to a centralized \u003cstrong\u003e\u003ccode\u003eState\u003c/code\u003e dictionary\u003c/strong\u003e, passing structured state rather than free-form text. Among mainstream frameworks, this is the most advanced in communication efficiency—it natively achieves what PACT aims for: \u003cstrong\u003ereplace natural-language chat with structured state.\u003c/strong\u003e But LangGraph\u0026rsquo;s limitation is that it requires developers to explicitly define the state schema, adding design overhead and making it less flexible for highly unstructured collaboration scenarios.\u003c/p\u003e\n\u003ch3 id=\"claude-code--codex-openai\"\u003e\n  Claude Code / Codex (OpenAI)\n  \u003ca class=\"heading-link\" href=\"#claude-code--codex-openai\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"Link to heading\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003eLink to heading\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h3\u003e\n\u003cp\u003eThe communication patterns of these two agentic coding tools are somewhat unique. They primarily interact with tools via protocols like MCP (Model Context Protocol), with relatively little internal agent-to-agent communication. More often, they adopt a \u0026ldquo;monolithic mega-context\u0026rdquo; strategy—dumping massive amounts of information into a long-context model at once and letting the model figure out understanding and scheduling on its own. This works at present (Gemini 3.1 Pro has a 2-million-token context window), but it\u0026rsquo;s unsustainable for the same fundamental reason: context windows may be getting larger, but attention density dilutes as the window grows.\u003c/p\u003e\n\u003chr\u003e\n\u003ch2 id=\"v-the-reasoning-model-era-the-thinking-trace-amplification-disaster\"\u003e\n  V. The Reasoning Model Era: The Thinking-Trace Amplification Disaster\n  \u003ca class=\"heading-link\" href=\"#v-the-reasoning-model-era-the-thinking-trace-amplification-disaster\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"Link to heading\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003eLink to heading\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h2\u003e\n\u003cp\u003eIf the problems discussed so far were merely \u0026ldquo;optimization targets,\u0026rdquo; the arrival of reasoning models turns them into \u0026ldquo;requirements.\u0026rdquo;\u003c/p\u003e\n\u003cp\u003eIn 2026, OpenAI GPT-5.5, DeepSeek-V4, Gemini 3.1 Pro Thinking, Claude Opus 4.8, and other reasoning models all default to (or optionally enable) long chain-of-thought reasoning internally. A single inference can generate thousands to tens of thousands of tokens of hidden thought:\u003c/p\u003e\n\u003cpre tabindex=\"0\"\u003e\u003ccode\u003e\u0026#34;Hmm, let me reconsider this problem. Starting from first principles...\u0026#34;\n\u0026#34;Wait, there\u0026#39;s a flaw in my earlier reasoning...\u0026#34;\n\u0026#34;Actually, there\u0026#39;s another possibility...\u0026#34;\n\u0026#34;Taking all of the above into account, my conclusion is...\u0026#34;\n\u003c/code\u003e\u003c/pre\u003e\u003cp\u003eThese thinking traces have real value—they make models reason deeper and more accurately. The problem: \u003cstrong\u003eif the MAS framework doesn\u0026rsquo;t isolate these traces, they get shoved wholesale into shared history.\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe consequences are catastrophic:\u003c/p\u003e\n\u003col\u003e\n\u003cli\u003e\u003cstrong\u003eToken billing spirals out of control\u003c/strong\u003e: a 2,000-token final output may sit atop 8,000 tokens of chain-of-thought. Three downstream agents each then replicate this 8,000-token context overhead.\u003c/li\u003e\n\u003cli\u003e\u003cstrong\u003eCross-agent contamination\u003c/strong\u003e: Agent B reads Agent A\u0026rsquo;s internal deliberation (\u0026ldquo;Hmm, wait\u0026hellip;\u0026rdquo; \u0026ldquo;Actually, there\u0026rsquo;s an issue\u0026hellip;\u0026rdquo;) and may be misled by A\u0026rsquo;s hesitation, making unnecessary corrective moves.\u003c/li\u003e\n\u003cli\u003e\u003cstrong\u003eAttention dilution\u003c/strong\u003e: research shows LLMs already struggle to process information in the middle of their context (\u0026ldquo;Lost in the Middle\u0026rdquo;). When the context is stuffed with other agents\u0026rsquo; thinking processes, the probability of critical information being buried soars.\u003c/li\u003e\n\u003c/ol\u003e\n\u003cp\u003e\u003cstrong\u003eThis is precisely where PACT and SupervisorAgent deliver their greatest value in the reasoning-model era: they enforce isolation of internal thinking traces, passing only final conclusions or structured action-state records between agents.\u003c/strong\u003e This isn\u0026rsquo;t a nice-to-have—it\u0026rsquo;s a prerequisite for reasoning models to participate in multi-agent collaboration.\u003c/p\u003e\n\u003chr\u003e\n\u003ch2 id=\"vi-the-industry-is-reinventing-tcp-a-brief-history-of-agent-communication-protocols\"\u003e\n  VI. The Industry Is \u0026ldquo;Reinventing TCP\u0026rdquo;: A Brief History of Agent Communication Protocols\n  \u003ca class=\"heading-link\" href=\"#vi-the-industry-is-reinventing-tcp-a-brief-history-of-agent-communication-protocols\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"Link to heading\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003eLink to heading\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h2\u003e\n\u003cp\u003eAcross 2025–2026, the industry has recognized that inter-agent communication needs standardization, not each framework rolling its own. Several major protocols are converging:\u003c/p\u003e\n\u003ctable\u003e\n  \u003cthead\u003e\n      \u003ctr\u003e\n          \u003cth\u003eProtocol\u003c/th\u003e\n          \u003cth\u003eSteward\u003c/th\u003e\n          \u003cth\u003ePurpose\u003c/th\u003e\n          \u003cth\u003eCommunication-Efficiency Relevance\u003c/th\u003e\n      \u003c/tr\u003e\n  \u003c/thead\u003e\n  \u003ctbody\u003e\n      \u003ctr\u003e\n          \u003ctd\u003e\u003cstrong\u003eMCP (Model Context Protocol)\u003c/strong\u003e\u003c/td\u003e\n          \u003ctd\u003eAnthropic\u003c/td\u003e\n          \u003ctd\u003eAgent ↔ tool connection\u003c/td\u003e\n          \u003ctd\u003eStandardizes tool invocation via JSON-RPC, dramatically reducing \u0026ldquo;explain how the tool works\u0026rdquo; prompt tokens\u003c/td\u003e\n      \u003c/tr\u003e\n      \u003ctr\u003e\n          \u003ctd\u003e\u003cstrong\u003eA2A (Agent-to-Agent)\u003c/strong\u003e\u003c/td\u003e\n          \u003ctd\u003eGoogle, donated to Linux Foundation\u003c/td\u003e\n          \u003ctd\u003eAgent ↔ Agent peer-to-peer orchestration\u003c/td\u003e\n          \u003ctd\u003eEnables capability exchange via Agent Cards; defines standards for \u0026ldquo;agents discovering other agents\u0026rdquo;\u003c/td\u003e\n      \u003c/tr\u003e\n      \u003ctr\u003e\n          \u003ctd\u003e\u003cstrong\u003eACP (Agent Communication Protocol)\u003c/strong\u003e\u003c/td\u003e\n          \u003ctd\u003eIBM\u003c/td\u003e\n          \u003ctd\u003eBroker-based message proxy\u003c/td\u003e\n          \u003ctd\u003eBroker pattern; reduces point-to-point complexity\u003c/td\u003e\n      \u003c/tr\u003e\n      \u003ctr\u003e\n          \u003ctd\u003e\u003cstrong\u003eANP (Agent Network Protocol)\u003c/strong\u003e\u003c/td\u003e\n          \u003ctd\u003eCommunity\u003c/td\u003e\n          \u003ctd\u003eDecentralized agent marketplace\u003c/td\u003e\n          \u003ctd\u003eTrusted agent discovery and capability negotiation\u003c/td\u003e\n      \u003c/tr\u003e\n  \u003c/tbody\u003e\n\u003c/table\u003e\n\u003cp\u003eBut these protocols currently solve \u0026ldquo;how to find agents\u0026rdquo; and \u0026ldquo;how to connect agents\u0026rdquo;—\u003cstrong\u003enot one of them addresses \u0026ldquo;what agents should say, and how much, once they\u0026rsquo;re connected.\u0026rdquo;\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThis is exactly where the PACT paper\u0026rsquo;s significance lies: it attempts to define the communication \u003cstrong\u003econtent layer\u003c/strong\u003e, not the transport or discovery layer. The analogy to TCP/IP\u0026rsquo;s four-layer model is apt:\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003eA2A/ACP/MCP solve the \u0026ldquo;transport layer\u0026rdquo; and \u0026ldquo;application layer\u0026rdquo;—how data moves, how agents discover each other\u003c/li\u003e\n\u003cli\u003ePACT attempts to solve the \u0026ldquo;presentation layer\u0026rdquo;—\u003cstrong\u003ewhat the data should look like\u003c/strong\u003e\u003c/li\u003e\n\u003c/ul\u003e\n\u003cp\u003eJust as HTTP/2\u0026rsquo;s introduction of header compression wasn\u0026rsquo;t a nice-to-have but the foundational technology that enabled the web to scale, agent communication content compression won\u0026rsquo;t be an optional optimization—it\u0026rsquo;s the prerequisite for growing from 5 agents to 500.\u003c/p\u003e\n\u003chr\u003e\n\u003ch2 id=\"vii-three-actionable-recommendations\"\u003e\n  VII. Three Actionable Recommendations\n  \u003ca class=\"heading-link\" href=\"#vii-three-actionable-recommendations\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"Link to heading\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003eLink to heading\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h2\u003e\n\u003cp\u003eIf you\u0026rsquo;re building or operating a multi-agent system, these three checks are immediately actionable:\u003c/p\u003e\n\u003ch3 id=\"1-audit-your-agent-communication-logs\"\u003e\n  1. Audit Your Agent Communication Logs\n  \u003ca class=\"heading-link\" href=\"#1-audit-your-agent-communication-logs\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"Link to heading\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003eLink to heading\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h3\u003e\n\u003cp\u003eGrab the full conversation history of any multi-agent task and count: how much content is redundant? How many tokens are utterly useless to downstream agents? How many \u0026ldquo;Thanks for your analysis\u0026rdquo; and \u0026ldquo;Got it, let me handle this\u0026rdquo; social niceties are burning tokens?\u003c/p\u003e\n\u003ch3 id=\"2-isolate-thinking-traces\"\u003e\n  2. Isolate Thinking Traces\n  \u003ca class=\"heading-link\" href=\"#2-isolate-thinking-traces\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"Link to heading\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003eLink to heading\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h3\u003e\n\u003cp\u003eIf your agents use reasoning models, ensure chain-of-thought is stripped before passing to downstream agents. Don\u0026rsquo;t feed \u0026ldquo;Hmm, let me think\u0026hellip;\u0026rdquo; to the next agent—it doesn\u0026rsquo;t need to know how you agonized.\u003c/p\u003e\n\u003ch3 id=\"3-replace-natural-language-chat-with-structured-state\"\u003e\n  3. Replace Natural-Language Chat with Structured State\n  \u003ca class=\"heading-link\" href=\"#3-replace-natural-language-chat-with-structured-state\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"Link to heading\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003eLink to heading\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h3\u003e\n\u003cp\u003eIf you\u0026rsquo;re using LangGraph, you\u0026rsquo;re already doing this—but ensure the state schema design converges (only contains what downstream needs) rather than expands (throwing everything in). If you\u0026rsquo;re using CrewAI or AutoGen, consider inserting a SupervisorAgent-like filtering layer between agents.\u003c/p\u003e\n\u003chr\u003e\n\u003ch2 id=\"coda\"\u003e\n  Coda\n  \u003ca class=\"heading-link\" href=\"#coda\"\u003e\n    \u003ci class=\"fa-solid fa-link\" aria-hidden=\"true\" title=\"Link to heading\"\u003e\u003c/i\u003e\n    \u003cspan class=\"sr-only\"\u003eLink to heading\u003c/span\u003e\n  \u003c/a\u003e\n\u003c/h2\u003e\n\u003cp\u003eThe question in 2023 was: \u0026ldquo;How do we get multiple agents to work together?\u0026rdquo;\u003c/p\u003e\n\u003cp\u003eIn 2024: \u0026ldquo;How do we get multiple agents to work together efficiently?\u0026rdquo;\u003c/p\u003e\n\u003cp\u003eIn 2025: \u0026ldquo;How do we get 1,000 agents to work together efficiently?\u0026rdquo;\u003c/p\u003e\n\u003cp\u003eIn 2026, the PACT and SupervisorAgent papers tell us: \u003cstrong\u003ethe answer lies not in better orchestration—but in less noise.\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eAgent communication protocols may be the next infrastructure-level opportunity. Just as TCP/IP defined how nodes on the internet handshake, \u003cstrong\u003ethe \u0026ldquo;presentation layer\u0026rdquo; of the Agent Communication Protocol will define how AI agents collaborate efficiently at scale.\u003c/strong\u003e And the PACT paper may be one of the earliest milestones on that path.\u003c/p\u003e\n\u003chr\u003e\n\u003cp\u003e\u003cstrong\u003eReferences\u003c/strong\u003e:\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003e\n\u003cp\u003eChen Huang et al., \u003cem\u003eWhat Should Agents Say? Action-state Communication for Efficient Multi-Agent Systems\u003c/em\u003e, arXiv: 2606.05304, 2026-06-06. SUTD.\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cem\u003eStop Wasting Your Tokens: Towards Efficient Runtime Multi-Agent Systems\u003c/em\u003e, ICLR 2026. SupervisorAgent.\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cem\u003eA Survey of Agent Interoperability Protocols: MCP, ACP, A2A, and ANP\u003c/em\u003e, arXiv: 2505.02279, 2025.\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eAnthropic, Claude Code Dynamic Workflows, 2026.\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eOpenAI, Codex Agent Architecture, 2026.\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003ePACT open-source implementation: \u003ca href=\"https://github.com/iNLP-Lab/PACT\"  class=\"external-link\" target=\"_blank\" rel=\"noopener\"\u003ehttps://github.com/iNLP-Lab/PACT\u003c/a\u003e\u003c/p\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n",
  "wordCount": 2303,
  "readingTime": 11,
  "tableOfContents": "\u003cnav id=\"TableOfContents\"\u003e\n  \u003cul\u003e\n    \u003cli\u003e\u003ca href=\"#i-the-physics-of-the-problem-the-on-communication-disaster\"\u003eI. The Physics of the Problem: The O(n²) Communication Disaster\u003c/a\u003e\u003c/li\u003e\n    \u003cli\u003e\u003ca href=\"#ii-what-the-pact-paper-found-agents-dont-need-to-write-essays\"\u003eII. What the PACT Paper Found: Agents Don\u0026rsquo;t Need to Write Essays\u003c/a\u003e\n      \u003cul\u003e\n        \u003cli\u003e\u003ca href=\"#21-five-communication-strategies-none-dominant\"\u003e2.1 Five Communication Strategies, None Dominant\u003c/a\u003e\u003c/li\u003e\n        \u003cli\u003e\u003ca href=\"#22-pact-turning-agent-communication-into-state-updates\"\u003e2.2 PACT: Turning Agent Communication into State Updates\u003c/a\u003e\u003c/li\u003e\n        \u003cli\u003e\u003ca href=\"#23-quantified-impact-not-incrementaltransformational\"\u003e2.3 Quantified Impact: Not Incremental—Transformational\u003c/a\u003e\u003c/li\u003e\n      \u003c/ul\u003e\n    \u003c/li\u003e\n    \u003cli\u003e\u003ca href=\"#iii-supervisoragent-a-different-route-same-destination\"\u003eIII. SupervisorAgent: A Different Route, Same Destination\u003c/a\u003e\u003c/li\u003e\n    \u003cli\u003e\u003ca href=\"#iv-how-current-mainstream-frameworks-handle-communication-a-reality-check\"\u003eIV. How Current Mainstream Frameworks Handle Communication: A Reality Check\u003c/a\u003e\n      \u003cul\u003e\n        \u003cli\u003e\u003ca href=\"#crewai\"\u003eCrewAI\u003c/a\u003e\u003c/li\u003e\n        \u003cli\u003e\u003ca href=\"#autogen-microsoft\"\u003eAutoGen (Microsoft)\u003c/a\u003e\u003c/li\u003e\n        \u003cli\u003e\u003ca href=\"#langgraph\"\u003eLangGraph\u003c/a\u003e\u003c/li\u003e\n        \u003cli\u003e\u003ca href=\"#claude-code--codex-openai\"\u003eClaude Code / Codex (OpenAI)\u003c/a\u003e\u003c/li\u003e\n      \u003c/ul\u003e\n    \u003c/li\u003e\n    \u003cli\u003e\u003ca href=\"#v-the-reasoning-model-era-the-thinking-trace-amplification-disaster\"\u003eV. The Reasoning Model Era: The Thinking-Trace Amplification Disaster\u003c/a\u003e\u003c/li\u003e\n    \u003cli\u003e\u003ca href=\"#vi-the-industry-is-reinventing-tcp-a-brief-history-of-agent-communication-protocols\"\u003eVI. The Industry Is \u0026ldquo;Reinventing TCP\u0026rdquo;: A Brief History of Agent Communication Protocols\u003c/a\u003e\u003c/li\u003e\n    \u003cli\u003e\u003ca href=\"#vii-three-actionable-recommendations\"\u003eVII. Three Actionable Recommendations\u003c/a\u003e\n      \u003cul\u003e\n        \u003cli\u003e\u003ca href=\"#1-audit-your-agent-communication-logs\"\u003e1. Audit Your Agent Communication Logs\u003c/a\u003e\u003c/li\u003e\n        \u003cli\u003e\u003ca href=\"#2-isolate-thinking-traces\"\u003e2. Isolate Thinking Traces\u003c/a\u003e\u003c/li\u003e\n        \u003cli\u003e\u003ca href=\"#3-replace-natural-language-chat-with-structured-state\"\u003e3. Replace Natural-Language Chat with Structured State\u003c/a\u003e\u003c/li\u003e\n      \u003c/ul\u003e\n    \u003c/li\u003e\n    \u003cli\u003e\u003ca href=\"#coda\"\u003eCoda\u003c/a\u003e\u003c/li\u003e\n  \u003c/ul\u003e\n\u003c/nav\u003e",
  "isDraft": false
}
