arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

大语言模型(LLM)可预测失败风险,但难以预测哪种协作协议能产生回报:推理任务中考虑成本的协议路由

LLMs Can Predict Failure Risk, But Struggle to Predict Which Collaboration Protocol Pays Off: Cost-Aware Protocol Routing Across Reasoning Tasks

Chih-Hsuan Yang, Jingyan Jiang, Cheng-Hau Yang, Vikram Vasudevan, Huihuo Zheng, Venkatram Vishwanath, Rajeev Thakur

arXiv 2608.14927首次发表:更新:

发表机构

Argonne National Laboratory; Oregon State University(阿贡国家实验室; 俄勒冈州立大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究探讨LLM多智能体系统中协作成本与收益的平衡,测试四种协作协议,发现LLM可预测失败风险但难以识别有效协议,置信度可支持初始协作升级,特定协议的成本感知路由仍待解决。

AI 中文摘要

多智能体大语言模型(LLM)系统可通过投入更多计算资源提升推理能力,但部署时需判断额外协作是否值得其成本。本研究通过在每个设置中固定求解器,在四种协议下运行所有问题来隔离该决策:直接求解(基线)、迭代自校正(Single)、规划器-执行器-评审员协作(PER)以及多智能体商议(Broadcast)。主要基准包含4181道竞赛级数学题;配对稳健性检验覆盖四个基准,涉及竞赛数学、生物学及更广泛科学领域,并采用两类求解器系列。在固定策略、训练型路由器和冻结LLM路由器中,保守策略会过度低估协作需求,而更高求解能力的冻结路由器则常过度高估协作需求。一种在答案后、协作前的gpt-oss-120b探针以0.8847的AUROC对基线失败案例进行排序(4151个可解析案例;95%置信区间[0.8732, 0.8955])。该分数对预测任何协作是否有帮助仍具参考性(0.7683的AUPRC),但对识别PER或Broadcast特有的价值则弱得多(AUPRC分别为0.1674和0.1041)。此外,答案前的自置信门在4.5万token时达到78.0%的求解率,而冻结gpt-oss-120b路由器在7.13万token时仅达到73.8%,回顾性固定顺序oracle则达到92.4%。在10组配对模型-条件设置中,oracle相比基线增加了23.2至58.3点的回顾性覆盖率,但协议特征因任务而异。在6组保留路由器评估的设置中,oracle差距仍为18.5至28.9点。因此,置信度可支持初始协作升级,但针对特定协议的成本感知路由仍未解决。

英文摘要

Multi-agent large language model (LLM) systems can improve reasoning by spending more computation, but deployment requires deciding when extra collaboration is worth its cost. We isolate this decision by running every problem under four protocols while holding the solver fixed within each setting: direct solving (Baseline), iterative self-correction (Single), planner-executor-reviewer collaboration (PER), and multi-agent deliberation (Broadcast). The primary benchmark comprises 4,181 competition-level math problems; paired robustness checks cover four benchmarks spanning competition math, biology, and broader science with two solver families. Across fixed policies, trained routers, and frozen LLM routers, conservative policies under-escalate, whereas higher-solve frozen routers often over-escalate. A post-answer, pre-collaboration gpt-oss-120b probe ranks Baseline failures with 0.8847 AUROC (4,151 parseable cases; 95% CI [0.8732, 0.8955]). The same score remains informative for predicting whether any collaboration helps (0.7683 AUPRC), but is much weaker for identifying PER- or Broadcast-specific value (0.1674 and 0.1041 AUPRC). Separately, the pre-answer self-confidence gate reaches 78.0% solve at 45K tokens, compared with 73.8% at 71.3K for a frozen gpt-oss-120b router and 92.4% for a retrospective fixed-order oracle. Across 10 paired model-condition settings, the oracle adds 23.2-58.3 points of retrospective coverage over Baseline, but protocol profiles vary by task. In the six settings with held-out router evaluations, oracle gaps remain 18.5-28.9 points. Confidence can therefore support initial escalation, while protocol-specific cost-aware routing remains unresolved.

Comments23 pages, 6 figures; includes appendices and ancillary aggregate-result CSV files

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑