AI 中文总结
研究多智能体大语言模型级联中的语义寄存器压缩现象,用三智能体管道量化压缩,通过五个架构变体分析驱动因素,发现压缩在不同领域强度不同且可泛化,提示级回归可解释大部分方差,对高风险领域安全评估有意义。
AI 中文摘要
多智能体大语言模型系统通常将复杂任务分解为专门角色。然而,这种模块化引入了一种表征风险:当中间智能体跨语言寄存器转换文本时,它们可能会系统地压缩准确下游决策所需的语义区别。我们将此现象称为语义寄存器压缩,并将其表征为多智能体级联中的一种可观察到的故障模式。使用三智能体管道(收集器 - 评估器 - 决策器),我们通过句子转换器嵌入空间中的标签间分离来量化压缩。在政治事实核查(LIAR)、情感分析(SST - 5)和医疗分诊(Triagegeist)中,关键评估在评估器阶段始终将标签可分离性降低41.7%,而身份直通几乎完全保留它。五个架构变体因果隔离了定向语义转换作为主要驱动因素。一个寻求可信度的变体产生最小的几何压缩,但将输出转向大多为真,表明转换价控制分布崩溃的方向独立于压缩幅度。压缩在三个领域以不同强度泛化:事实核查中为41.7%,情感分析中为27.2%,分诊中为20.0%。提示级回归解释了78%的方差,操作约束与较低压缩相关。这些结果表明语义寄存器压缩在多智能体大语言模型系统中是一种可测量和可泛化的现象,对高风险领域的安全评估有影响。
英文摘要
Multi-agent LLM systems commonly decompose complex tasks into specialized roles. However, this modularity introduces a representational risk: when intermediate agents transform text across linguistic registers, they can systematically compress the semantic distinctions needed for accurate downstream decisions. We term this phenomenon semantic register compression and characterize it as an observable failure mode in multi-agent cascades. Using a three-agent pipeline (Collector-Evaluator-Decider), we quantify compression via inter-label separation in sentence-transformer embedding space. Across political fact-checking (LIAR), sentiment analysis (SST-5), and medical triage (Triagegeist), critical evaluation reduces label separability at the Evaluator stage, while identity passthrough preserves it nearly fully. Five controlled variants show that geometric change depends on the specific intermediate transformation rather than on the mere presence of an additional cascade stage. A credibility-seeking variant expands rather than compresses inter-label separation, while shifting outputs toward mostly-true, demonstrating that transformation valence controls both the direction and the sign of geometric change independently of compression magnitude. Compression generalizes across the three domains with domain-dependent intensity (10.3% in fact-checking, 28.2% in sentiment, 9.1% in triage). A 20-level prompt gradient reveals a non-monotonic compression profile: balanced evaluative prompts produce the strongest compression, while extreme critical prompts show irregular moderate compression. These results demonstrate that semantic register compression is a measurable and generalizable phenomenon in multi-agent LLM systems, with implications for safety evaluation in high-stakes domains.
Comments16 pages, 2 figures, 4 tables. Revised version with corrected English-prompt experiments and updated cross-domain and prompt-gradient results