arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

可助推性:推理模型遵循置信信号而不追踪自身能力

Nudgeability: Reasoning Models Follow Confidence Signals Without Tracking Their Own Competence

Rohit Saxena, Utkarsh Upadhyay

arXiv 2609.34572首次发表:更新:

AI 中文总结

本研究提出可助推性指标,通过插入信心或怀疑句子测试推理模型的委派决策,发现模型对置信信号敏感但针对性差,仅比随机基线高2个百分点。

AI 中文摘要

能够调用工具的推理语言模型在推理过程中必须决定是独立作答还是委派给工具。任何用于此决策的自我反思机制都必须回答三个问题:反思信号来自何处(口头报告、输出分布、隐藏状态、独立预测器),如何呈现给模型(数值预测、置信令牌、提示注入),以及它是否改变模型后续行动。我们隔离第三个问题。在原本相同的推理轨迹中的固定点,我们插入一个表达信心或怀疑的第一人称句子;模型随后继续推理并选择直接回答还是调用工具。比较这些反事实延续可衡量反思信号对委派的因果效应。我们将这种行为反应称为可助推性,并沿两个维度衡量:敏感性,即信心和怀疑改变委派率的强度;以及针对性,即委派是否对模型无法独立解决的问题增加,而对模型能够解决的问题减少。在来自三个家族(Qwen、Gemma和GLM)的九个中小型开放权重推理模型和两个任务中,模型始终敏感:怀疑增加委派,信心减少委派,信心到怀疑的中位摆动为20.6个百分点,而较大的提供商服务模型为53至70个百分点。这种响应性针对性较差:中位数42%的诱导翻转是良好针对的,仅比随机选择基线高2个百分点。因此,信心语言是委派的强控制面,但当前模型仅根据其实际能力弱使用它。可助推性提供了一种简单的、无需训练后处理的方法来评估敏感性和针对性,随着内生自我反思机制的成熟。

英文摘要

Reasoning language models that can call tools must decide during inference whether to answer unaided or delegate. Any self-reflection mechanism for this must answer three questions: where the reflective signal comes from (verbal reports, output distributions, hidden states, a separate predictor), how it is presented to the model (numerical prediction, confidence token, prompt injection), and whether it changes the model's subsequent action. We isolate the third question. At a fixed point in otherwise identical reasoning trajectories, we insert a single first-person sentence expressing either confidence or doubt; the model then continues reasoning and chooses whether to answer directly or call a tool. Comparing these counterfactual continuations measures the causal effect of the reflective signal on delegation. We call this behavioral response Nudgeability and measure it along two dimensions: sensitivity, how strongly confidence and doubt change delegation rates, and targeting, whether delegation increases for problems the model cannot solve unaided and decreases for those it can. Across nine small-to-medium open-weight reasoning models from three families (Qwen, Gemma, and GLM) and two tasks, models are consistently sensitive: doubt increases delegation and confidence decreases it, with a median confidence-to-doubt swing of 20.6 percentage points, and 53 to 70 points for the larger provider-served models. This responsiveness is poorly targeted: a median 42% of induced flips are well-targeted, only a +2 percentage-point lift over a random-selection baseline. Confidence language is thus a strong control surface for delegation, but current models use it only weakly in accordance with their actual competence. Nudgeability offers a simple, post-training-free way to evaluate both sensitivity and targeting as endogenous self-reflection mechanisms mature.

Comments22 pages, 5 figures, 9 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑