通信瓶颈:语言模型中树结构表达式序列化的往返研究
The Communication Bottleneck: A Round-Trip Study of Tree-Structured Expression Serialization in Language Models
浏览论文内容
中文总结 AI 辅助
本研究通过往返协议实证发现,语言模型间以自然语言传递树结构表达式存在有损且不对称的瓶颈,生成阶段是主要失败源,且可通过共享结构的微调数据有效缓解。
中文摘要 AI 辅助
当语言模型以思维链方式推理或交换自由文本中间结果时,它们将结构化信息序列化为自然语言。有多少树结构的组合内容能够在这种瓶颈中幸存下来?我们提出了一种往返协议,以实证方式回答树结构表达式的这一问题。生成器将程序化生成的算术表达式转换为文字题,独立的提取器仅从文字题中恢复表达式,符号等价性提供了精确的判定标准。评估十六个模型的所有成对组合,得到一个通信矩阵,其边缘分布将生成质量与提取质量分离。三个主要发现随之浮现。第一,该信道是有损且不对称的:交换哪个模型生成和哪个模型提取会使准确率变化高达60.4个百分点,而最佳配对通过在不同端组合不同模型(而非两端使用同一模型)达到92.9%的准确率。第二,至少73.6%的往返失败源于生成阶段,且难度由树结构(运算符数量、深度、右分支)驱动,而非模型家族。第三,该信道是可训练的:约3600个与评估共享运算符和树形状的微调示例,使每个开放权重模型都超过未训练的Gemini-3.1-Pro,后者是在匹配语义下的上界。一个使用新运算符和词汇的不相交域设置也提升了每个开放权重模型,确认增益并非匹配语义的伪影,尽管与前沿模型仍有差距。这些结果共同表明,当模型通过自然语言交流层级结构时,树结构表达式序列化是主要的限制因素。
英文摘要
When language models reason in chain-of-thought or exchange free-text intermediates, they serialize structured information into natural language. How much tree-structured compositional content survives this bottleneck? We propose a round-trip protocol that answers this question empirically for tree-structured expressions. A generator converts a procedurally generated arithmetic expression into a word problem, a separate extractor recovers the expression from the word problem alone, and symbolic equivalence provides an exact oracle. Evaluating all pairwise combinations of sixteen models yields a communication matrix whose marginals separate generation quality from extraction quality. Three main findings emerge. First, the channel is lossy and asymmetric: swapping which model generates and which extracts shifts accuracy by up to 60.4 points, and the best pair reaches 92.9% by combining different models on each end rather than the same model on both. Second, at least 73.6% of round-trip failures originate at generation, and difficulty is driven by tree structure (operator count, depth, right-branching) rather than model family. Third, the channel is trainable: ~3600 fine-tuning examples that share the evaluation's operators and tree shapes lift every open-weight model above untrained Gemini-3.1-Pro, an upper bound under matched semantics. A disjoint-domain regime with new operators and vocabulary also raises every open-weight model, confirming the gain is not an artifact of matched semantics, though a gap to the frontier remains. Together these results identify tree-structured expression serialization as a primary limiting factor when models communicate hierarchical structure through natural language.
发表机构
- Apple(苹果公司)
机构由 AI 辅助整理,请以论文原文为准。