发表机构
University of Southern California(南加州大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究发现,简短自然上下文可使Jev等决策模型在61.4%的初始正确决策上翻转至高置信度错误选项,暴露了当前决策模型的显著脆弱性。
AI 中文摘要
诸如Jev之类的专用决策模型将非结构化语言映射到有限选项上的概率分布,从而使其输出能够直接路由请求、选择工具并触发操作。然而,现实世界的输入很少孤立出现:它们会附带背景细节和周围上下文。我们发现,能够自然融入此上下文的简短添加内容,即便正确答案保持不变,也可能使原本正确的决策发生偏转。为研究这一行为,我们为每个初始正确的条目固定一个错误的备选选项,并利用模型的选项概率来优化流畅的上下文添加内容,同时保留来源、问题、选项和黄金答案。在64次被接受的靶向评估中,优化器识别出的上下文使Jev在508个初始正确决策中的312个(61.4%)发生翻转;在229个案例中,Jev对固定错误选项分配了至少0.7的概率。在七个数据集上,另外三个决策系统在其初始回答正确的决策上显示出64.9%-73.2%的靶向翻转率。综合来看,这些结果暴露了当前决策模型的显著脆弱性:简短、看似普通的上下文可以将正确选择转变为高置信度的错误选择。由于这些模型直接将语言转化为下游选择,这种敏感性引发了对将其概率输出视为可靠决策接口的担忧。
英文摘要
An ordinary-looking background detail can turn a correct model decision into a confident mistake. We demonstrate this fragility in four decision systems, including Jev, across seven datasets covering knowledge, reasoning, and tool routing. Within 64 accepted target evaluations per decision, we uncover short context additions that redirect 61.4%-73.2% of each system's initially correct decisions toward a wrong option fixed in advance. The additions supply background or procedural information rather than explicit answer-selection instructions, leaving the original question and choices intact. We construct them through probability-guided context optimization, which uses shifts in the option distribution to refine surrounding text under naturalness and answer-preservation constraints. Redirection affects initially confident decisions, often produces high-confidence wrong choices, and transfers across models. In blinded human evaluation, 91.6% of 250 sampled successful contexts are judged natural, answer-preserving, and free of decisive answer-changing evidence by a majority of three independent annotators. These findings expose a weakness in current decision models: context that looks entirely compatible with an input can redirect the choices that agents, routers, and evaluators rely on.
Comments46 pages, 8 figures, 30 tables. V2: substantially strengthened empirical evaluation with blinded human validation, a budget-matched independent-generation baseline, and repeatability analyses; expanded related work and updated author list. Homepage: https://xzx34.github.io/jevout/ ; Code: https://github.com/xzx34/JevOut