Flat Score, Amplified Failures:误差预算如何掩盖量化大语言模型智能体中的损害
Flat Score, Amplified Failures: How the Error Budget Masks Damage in Quantized LLM Agents
浏览论文内容
中文总结 AI 辅助
本文以τ²-bench为基准,发现量化大语言模型智能体的分数看似平稳,但会放大原有工具调用失败,其损害被基准的10个错误预算掩盖,缩小预算可暴露该问题,相关诊断方法可与任务奖励一同报告。
中文摘要 AI 辅助
4位权重的后训练量化被广泛报道几乎无损失,本文针对多轮工具调用智能体(该场景下量化影响最显著)验证这一结论。在τ²-bench基准上,针对密集型和混合专家(MoE)两类开放权重模型家族、两个领域(共8个单元,每个单元456个回合,权重精度分别为16位、8位、4位),量化在标准指标上看似无损失:无单元的分数变化经多重比较校正后仍显著,且在流程损害最大的单元中,等价测试将变化限制在±7.5分以内。但流程层面呈现不同结果:量化会放大模型全精度时已存在的失败(电信领域的工具名称幻觉、零售领域的实体错误,趋势一致),失败量最多增至2.5倍(每个任务增加17.6分),且未产生本质上的新失败;各精度下的失败集合相同(秩相关系数≥0.94,仅0.18%为新事件)。分数保持平稳是因为基准的10个错误预算吸收了额外失败;将预算缩小至2个错误时,仅在量化增加错误量的那个单元重新显现17分的分数差距,完全符合掩盖机制的预测。对5个电信领域模型在各精度下运行针对性错误修复提示,可精准消除该损害。本文提出的两种诊断方法(逐通道错误率、缩小预算下的成功率)均来自基准已收集的日志,建议将其与任务奖励一同报告。
英文摘要
Post-training quantization to 4-bit weights is widely reported to be nearly lossless. We test this claim for multi-turn, tool-calling agents, where it now matters most. On $τ^2$-bench, across two open-weight model families in dense and MoE variants and two domains (eight cells, 456 episodes each, at 16-, 8-, and 4-bit weights), quantization indeed looks free on the standard metric. No cell shows a score change that survives multiple-comparison correction, and in the cell that carries the largest process damage, equivalence testing bounds the change within $\pm$7.5 points. The process tells a different story. Quantization amplifies the failure the model already exhibits at full precision (tool-name hallucination in telecom, with the same directional trend in retail entity errors) by up to 2.5$\times$ in volume (+17.6 points per task), while creating essentially no new failures. The failure set is the same at every precision (rank correlation $\geq$ 0.94, 0.18% novel events). The score stays flat because the benchmark's ten-error budget absorbs the extra failures. Shrinking the budget to two errors re-exposes a score gap of 17 points, and it does so only in the one cell where quantization added error volume, exactly as the masking account predicts. A targeted error-repair prompt, run for five telecom models at every precision, removes the damage exactly and only where it lives. Both diagnostics, the per-channel error rate and success under a shrinking budget, come from logs benchmarks already collect; we suggest reporting them alongside task reward.