发表机构
The Ohio State University(俄亥俄州立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究通过因果审计多智能体LLM的中继KV缓存,发现仅当接收方需要发送方私有信息时,潜在通信才会发挥作用,且不匹配缓存审计是验证潜在想法传输的必要手段。
AI 中文摘要
多智能体大语言模型系统会中继键值(KV)缓存而非文本,并将其性能提升归因于交换的“潜在想法”。这一归因是关于中继的是哪个示例的缓存,而非仅仅是中继了某个缓存。我们在已发布的系统中对其进行因果审计。我们将缓存替换为错乱(示例不匹配)、置零和矩匹配的随机对应物,设置两种场景:接收方是否需要发送方的私有信息。在接收方需要私有信息的场景中,性能达到上限:主干模型上与答案无关的中继仅为23%-25%,而该场景下的性能为100%,这一对比在三个模型家族、五个检查点以及一个散文文档问答表面均得到复现。在接收方不需要私有信息的场景中,预注册的五种子实验协议证实,在GSM8K和ARC-Challenge数据集上针对三个Qwen3规模模型,以及在8B规模的MedQA数据集上,通过Holm校正的TOST检验,性能差异在2.8个百分点内(其中一个单元在该范围内显示出微小的检测到的优势;第二个模型家族未检测到优势)。大的缓存效应不一定是配对效应:在一个自然单元中,置零中继会导致14.7个百分点的性能损失,而不匹配的缓存仅导致0.4个百分点的损失。此外,需求也不充分:在相同测试下,所测试的通道涵盖上限(LatentMAS的原生中继)、部分(KVComm的层子集)以及未检测到示例特定传输(C2C的已发布投影器)。基准差异本身无法确立潜在想法的传输;要确立这一点需要进行不匹配缓存审计,我们已发布该审计。
英文摘要
Multi-agent LLM systems relay key-value caches instead of text and credit their gains to exchanged "latent thoughts". That credit is a claim about which example's cache is relayed, not merely that one is. We audit it causally in released systems. The cache is replaced with deranged (mismatched-example), zeroed, and moment-matched random counterparts, under two regimes defined by whether the receiver needs the sender's private information. Where it does, the battery reads ceiling: 100% against 23-25% for answer-irrelevant relays on the primary backbone, a contrast replicated across three families, five checkpoints, and a prose document-QA surface. Where it does not, a pre-registered five-seed protocol establishes equivalence within 2.8 points, a margin anchored to the audited system's reported gain, under Holm-corrected TOST on GSM8K and ARC-Challenge across three Qwen3 scales and on MedQA at 8B (one cell shows a small detected advantage inside the margin); a second family shows no detected advantage. A large cache effect need not be a pairing effect. In one natural cell, zeroing the relay costs 14.7 points; a mismatched cache, 0.4. Nor is need sufficient: under the same test, delivered channels span ceiling (LatentMAS's native relay), partial (KVComm's layer subset), and no detected example-specific transfer (C2C's released projector). Benchmark deltas do not by themselves establish latent-thought transmission; establishing it takes a mismatched-cache audit, which we release.
Commentsv2: metadata only, fixed abstract formatting