发表机构
Duke University(杜克大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究量化对语言模型智能体工具故障恢复的影响,发现8位与4位变体的恢复比较因提示词和评估设计而改变方向,强调评估需匹配任务、报告全流程成功率并量化不确定性。
AI 中文摘要
训练后量化降低了部署语言模型智能体的成本,但其对临时工具故障恢复的影响可能取决于恢复的评估方式。我们在二十个确定性工具使用任务和五个提示词上比较了Llama-3.1-8B-Instruct和Qwen2.5-7B-Instruct的8位与4位变体。8位与4位的恢复比较在提示词和评估目标之间会改变方向。在相同提示词下两个变体均无故障完成的任务上,Llama的差异范围为0到+20.2个百分点,Qwen的差异范围为-50.0到+35.0个百分点。全流程点估计在所有五个提示词下均偏向8位Llama,而Qwen的比较在提示词之间改变方向。评估目标也可能逆转结果。对于Llama在某个提示词下,仅对每个变体自身干净通过的任务进行评分,4位变体领先17.5个百分点;对两个变体的相同任务进行评分则无差异,而对全流程评分则8位变体领先28.3个百分点。执行器宽容度是第三个此类选择。使用严格的输出解析对相同日志重新评分(在该提示词下8位Llama的违规频率远高于4位Llama),将+28.3变为-15.0,而Qwen基本保持不变。这些发现表明,单个提示词、单个筛选任务集和单个评分策略无法确立关于量化智能体鲁棒性的稳定结论。评估应在匹配的任务上比较变体,为部署决策报告全流程成功率,说明评分策略,并量化跨任务的不确定性,而非注入故障位置。
英文摘要
Post-training quantization reduces the cost of deploying language-model agents, but its effect on recovery from temporary tool failures can depend on how recovery is evaluated. We compare 8-bit and 4-bit variants of Llama-3.1-8B-Instruct and Qwen2.5-7B-Instruct on twenty deterministic tool-use tasks and five prompts. The 8-bit-4-bit recovery comparison changes direction across prompts and evaluation targets. On tasks that both variants complete without faults under the same prompt, the difference ranges from 0 to +20.2 percentage points for Llama and from -50.0 to +35.0 points for Qwen. Full-pipeline point estimates favor 8-bit Llama under all five prompts, whereas the Qwen comparison changes direction across prompts. The evaluation target can also reverse the result. For Llama under one prompt, scoring each variant only on its own clean-passing tasks favors 4-bit by 17.5 points; scoring the same tasks for both variants gives no difference, while scoring the full pipeline favors 8-bit by 28.3 points. Executor leniency is a third such choice. Rescoring the same logs with strict output parsing, which 8-bit Llama violates far more often than 4-bit Llama under that prompt, turns that +28.3 into -15.0 while leaving Qwen essentially unchanged. These findings show that one prompt, one screened task set, and one scoring policy do not establish a stable conclusion about quantized-agent robustness. Evaluations should compare variants on matched tasks, report full-pipeline success for deployment decisions, state the scoring policy, and quantify uncertainty across tasks rather than injected fault sites.
CommentsAccepted at the NeurIPS 2026 Workshop on Small Language Models for Agentic Systems (SLM-Agents). 7 pages, 2 figures, 2 tables, plus appendix