arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

选择过多:研究大规模下LLM的决策制定

Overwhelmed by Choice: Studying LLM Decision Making at Scale

Yu-Chi Lin, Aryan Seth, Anshul Aravind, Eugene Lee, Tanmay Parekh, Nanyun Peng, Kai-Wei Chang

arXiv 2609.32809首次发表:更新:

AI 中文总结

本研究系统评估了候选集规模对LLM决策的影响,发现准确率随候选数增加而下降,归因于金标边际崩溃和早期偏好难以逆转,并提出层次划分和排列推断方法,在N=160时提升约20个百分点。

AI 中文摘要

多项选择和候选者选择评估被广泛用于评估LLM的推理和决策能力,然而大多数基准测试包含相对较小的候选集。目前尚不清楚在这些设置下得出的结论在候选空间扩大时是否仍然有效。我们系统性地评估了随着竞争候选者数量增加时LLM的表现,发现在不同任务、提示策略和模型规模上,准确率均出现显著下降。受控分析表明,标准的长上下文检索解释无法完全解释这种下降。相反,我们识别出两种系统性的失败模式。第一,金标边际崩溃:正确答案与最强干扰项之间的分数差距逐渐缩小,这主要是由于对正确答案的置信度减弱所致。第二,早期候选者的偏好越来越难以被推翻,而后期候选者对最终预测的影响越来越弱。基于这些发现,我们评估了层次划分和基于排列的推断方法,在HotpotQA和MIMIC上,当N=160时,准确率提高了约20个百分点。总体而言,我们的结果将候选集规模确定为评估协议中的一个重要变量,并表明在小型选项上的强表现并不一定意味着在大型候选比较中具有稳健性。

英文摘要

Multiple-choice and candidate-selection evaluations are widely used to assess LLM reasoning and decision-making, yet most benchmarks contain relatively small candidate sets. It remains unclear whether conclusions drawn from these settings remain valid as the candidate space scales. We systematically evaluate LLMs as the number of competing candidates increases and find substantial accuracy degradation across tasks, prompting strategies, and model scales. Controlled analyses show that standard long-context retrieval explanations cannot fully account for this degradation. Instead, we identify two systematic failure patterns. First, gold-margin collapse: the score gap between the correct answer and the strongest distractor progressively shrinks, driven primarily by weakening confidence in the correct answer. Second, earlier candidate preferences become increasingly difficult to overturn, with later candidates exerting progressively weaker influence on the final prediction. Motivated by these findings, we evaluate hierarchical partitioning and permutation-based inference, which improve accuracy by roughly 20 percentage points at $N=160$ on both HotpotQA and MIMIC. Overall, our results identify candidate-set scale as an important evaluation-protocol variable and show that strong small-option performance does not necessarily imply robust large-scale candidate comparison.

CommentsAccepted at TAE (Trust-AI-Eval): Can We Trust AI Evaluation?, NeurIPS 2026 Workshop. 23 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑