AI 中文总结
研究多跳智能体中继中消息格式效应,引入可控测试平台,发现格式效应取决于层级,强中继近乎无损,弱中继下跨格式差异增大,配对分叉注入中错误值保留情况与真实值保留匹配,结构提供错误定位通道,格式选择应依最弱中继。
AI 中文摘要
当语言模型智能体相互传递信息时,消息格式重要吗?两种文献观点不一:格式优化研究表明结构化消息能降低成本且不影响准确性,而格式限制研究发现施加结构会降低生成质量,且两者都未考量消息多跳传递时的情况,此时复制保真度而非一次性生成起主导作用。我们引入了一个可控的中继测试平台:十二个通过编程生成的原子事实摘要以五种格式(自由自然语言、精确指令自然语言、JSON、三元组、键值对)逐跳重新编码,经过六跳传递,由固定的强大评分器根据编程的地面真值进行评分,涵盖两个中继能力层级、一种认知负荷条件和一种配对分叉错误注入。我们发现消息格式效应取决于层级。(i)在忠实中继指令下,强大的中继几乎无损——未出现记录中的“传话游戏”崩溃情况——增加每跳认知负荷时,格式级保真度不变(在±1.8分范围内),同时生成成本提高24 - 53%。(ii)在弱(15亿参数)中继下,六跳召回率的跨格式差异增长了8.7倍(从2.3分增至20.5分),由两种相反机制驱动——刚性格式支付的编码代价以及特定于固定键JSON模式的抗漂移性——这在传递过程中翻转了格式排名。(iii)在配对分叉注入中,注入的错误值一旦出现,在每种格式的83 - 100%的链中都会持续到最后一跳,与每种格式对真实值的保留情况紧密匹配,对相邻事实没有可检测到的附带损害。结构提供了一个忠实、错误定位的通道——而非纠错码——格式选择应遵循管道中最弱的中继。
英文摘要
When LLM agents hand information to one another, does the message format matter? Two literatures disagree: format-optimization work reports that structured messages cut cost without hurting accuracy, while format-restriction studies find that imposing structure degrades generation. Neither line has measured what happens when messages traverse multiple hops, where copy fidelity, rather than one-shot generation quality, dominates. We introduce a controlled relay testbed in which briefs of twelve programmatic atomic facts are re-encoded hop by hop in five formats (free natural language, precision-instructed NL, JSON, triples, key-value) over six hops, scored against programmatic ground truth by a fixed strong grader, across two relay-capability tiers, a cognitive-load condition, and a paired-fork error injection. We find that (i) a strong relay is nearly lossless for every format (hop-6 QA recall $\geq 0.973$), with residual loss concentrated at the first encoding step; (ii) per-hop cognitive load raises generation cost by 24-53% while fidelity changes stay within $\pm 1.8$ points; (iii) under a weak 1.5B relay, the across-format dispersion of hop-6 recall grows by a factor of $8.7$ (CI 5.3-15.5), driven by an encode-drift trade-off that flips the format ranking in transit; and (iv) once an injected error is present, every format propagates it faithfully (surface persistence 83-100%) and no format cascades collateral damage onto neighboring facts. Structure buys a faithful, error-localizing channel, not an error-correcting code.