发表机构
Constructor University(康斯坦丁大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究大语言模型智能体间潜在通信,构建三个通信通道,通过稀疏自编码器分析信息损失,经实验发现文本通信会丢失信息,潜在通道在跨语言概念任务上与文本通道匹配但未超,得出丢失特征多编码表面形式的结论,还指出明确潜在通信优势需更深入任务及分析框架。
AI 中文摘要
多智能体系统在诸多场景和职业中被广泛应用,通常依靠明文消息传递进行智能体间通信。我们推测,当需要传达复杂概念时,大语言模型可能拥有超越文本表达能力的世界模型。为此,我们通过结构化实验来验证这一假设。研究表明,通过文本通信的语言模型智能体会丢失信息,我们借助稀疏自编码器(SAE)特征分析对其进行量化。我们构建了三个通信通道并测量每个通道中的概念区分信息。结果显示,SAE稀疏通道在28倍压缩时保持99.4%的探测准确率,而文本通道仅为80.4%。对于跨架构通信,使用稀疏潜在空间对齐时,Procrustes对齐在Llama和Mistral之间实现92%的top-1检索率。通过文本往返进行特征生存分析发现,文本序列化会破坏88%的SAE特征。我们将损失归因于身份替换而非衰减。在任务级评估中,潜在通道在跨语言概念任务上与文本通道匹配但未超越。用潜在特征进行文本增强并无益处。由此得出结论:丢失的特征大多或完全编码表面形式而非与任务相关的语义。要明确潜在通信相对于文本通道的实际优势,需要更深入的任务及相应分析框架。
英文摘要
Multi-agent systems (MAS) are utilized in many contexts and many professions. Those MAS rely on inter-agent communication, usually implemented by clear-text message passing. We hypothesize that Large Language Models may have a world model at their disposal that exceeds expressibility in text when complex concepts need to be communicated. Our aim is to approach a proof of this hypothesis with structured experiments. In this work, we show that LLM agents communicating via text lose information, which we quantify via Sparse Autoencoder (SAE) feature analysis. We construct three communication channels and measure concept-discriminating information in each. We first show that the SAE-sparse channel retains a 99.4% probe accuracy at 28-fold compression over the dense-latent channel vs 80.4% for the text channel. We then proceed to examine the same for cross-architecture communication by using sparse latent space alignment. We find for Procrustes alignment a 92% top-1 retrieval between Llama and Mistral. Using a text round-trip, we perform feature survival analysis to find that text serialization destroys 88% of SAE features, replacing them with a different feature set. We attribute the loss to identity replacement, not attenuation. By our analysis, we were able to attribute a 3-10pp performance penalty to the linear Procrustes alignment, improving with nonlinear alignment methods. In a task-level evaluation we find that the latent channel matches the text channel on cross-lingual concept tasks but never exceeds it. Text augmentation with latent features provides no benefit, leading us to negative conclusions for the initial hypothesis: lost features mostly or completely encode surface form, not task-relevant semantics. To pinpoint the practical advantage of latent communication over a text channel, deeper tasks eliciting complex concepts and an corresponding analysis framework are needed.