分解税:LLM流水线在其自身接口处损失高达40个准确率点
The Decomposition Tax: LLM Pipelines Lose Up to 40 Accuracy Points at Their Own Interfaces
浏览论文内容
中文总结 AI 辅助
本研究揭示LLM流水线在阶段接口处因信息丢失而损失高达40.5个准确率点,提出在损失后重新接地并保持数量间关系来降低分解税。
中文摘要 AI 辅助
一个四阶段的LLM流水线在其自身接口处损失高达40.5个准确率点(gemma-3-12B在MATH-500上,Holm校正p = 1.66e-19;主要系列中最大的税)。我们保持模型、问题、阶段、阶段提示和完成预算固定,仅改变每个阶段是否仍能看到原始问题,并将准确率差异称为分解税。在来自九个组织的21个开放权重模型上,在GSM-Hard和MATH-500上,每个单元格n = 200个配对项目,118个主要系列测试中有70个通过了Benjamini-Hochberg校正,54个通过了Holm校正。在GSM-Hard上,安慰剂没有恢复任何东西:它携带至少60%的额外令牌和至多一个单词的问题。构建者一次设计一个阶段的流水线,而其账单在阶段之间的接口处到达。重写一个阶段的指令将gemma-3-12B的税从4.5点移动到36.5点,并且在一个列出数值数量的阶段中添加“它们之间陈述的每个关系”在MATH-500上降低了9个模型中的9个的税。重新接地,即向阶段再次显示原始问题,应放在损失之后。对于一个有损接口,在损失之后重新接地该阶段在7个模型中的7个上优于在其之前重新接地,在两个基准上都是如此;在MATH-500上,较早的修复在7个模型中的7个上比没有修复更差。较新的模型仍然支付:gemma-4-12B放弃37.0点,并且修复在我们测试的三个最新模型上全部成立。一个密封的保留测试反驳了我们注册的一个更强的规则,该规则根据接口和接收者类型预测支付阶段,因此我们通过一次测量一个阶段来定位税。处方有两个部分:在有损接口之后重新接地该阶段,并且如果一个阶段必须列出数量,告诉它保持关系。
英文摘要
A four-stage LLM pipeline gives up as much as 40.5 accuracy points at its own interfaces (gemma-3-12B on MATH-500, Holm-corrected p = 1.66e-19; the largest tax in the primary family). We hold model, problem, stages, stage prompts and completion budget fixed, vary only whether each stage can still see the original problem, and call the accuracy difference the decomposition tax. Across 21 open-weight models from nine organisations, on GSM-Hard and MATH-500 at n = 200 paired items per cell, 70 of 118 primary-family tests survive Benjamini-Hochberg correction and 54 survive Holm. On GSM-Hard, a placebo recovers nothing: it carries at least 60% of the extra tokens and at most one word of the problem. Builders design a pipeline one stage at a time, and its bill arrives at the interfaces between stages. Rewriting one stage's instruction moves gemma-3-12B's tax from 4.5 to 36.5 points, and adding "every relationship stated between them" to a stage that lists the numerical quantities lowers the tax on 9 of 9 models on MATH-500. Re-grounding, which shows a stage the original problem again, belongs after the loss. With one lossy interface, re-grounding the stage after it beats re-grounding the stage before it on 7 of 7 models on both benchmarks; on MATH-500 the earlier repair is worse than none on 7 of 7. Newer models still pay: gemma-4-12B gives up 37.0 points, and the repair holds on all three of the newest models we test. A sealed held-out test refuted a stronger rule we registered, which predicted the paying stage from the interface and receiver types, so we locate the tax by measuring one stage at a time. The prescription has two parts: re-ground the stage after the lossy interface, and if a stage must list the quantities, tell it to keep the relationships.