arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.17686cs.LGcs.AIcs.CL

缺失的“我不知道”:为什么三个推理可靠性发现汇聚于校准弃权(不执行)

The Missing "I Don't Know": Why Three Reasoning-Reliability Findings Converge on Calibrated Abstention

Srijith Ravikumar

首次发表
浏览论文内容

中文总结 AI 辅助

三项独立研究发现大语言模型可靠性问题均指向缺失的校准弃权(不执行)能力,提出通过三重评分、弃权率报告等评估改革来弥合理论差距。

中文摘要 AI 辅助

三项近期结果描述了看似互不相关的大语言模型可靠性问题。Yin等人(2026)表明推理强化学习会破坏工具可靠性表征。Suleymanov等人(2026)表明在安全约束生成下,大模型会重写被标记的片段,而小模型则截断。Bastounis等人(2024)证明,任何没有隐式“我不知道”函数的一致推理系统,在广泛的问题类别上必然无限频繁地产生幻觉。我们认为这些发现汇聚于单一干预措施:校准弃权(不执行)正是每项研究独立识别出的缺失能力,尽管它们所记录的不可用性——能力差距、策略差距和递归论差距——在每种情况下来源不同。诚实性后训练已缩小了已部署模型中的差距,但对Bastounis所识别类别的原则性闭合需要一个校准弃权(不执行)函数,而该函数在排行榜层面的训练信号缺失:主流基准对拒绝回答赋予零奖励,因此能够选择该函数的排行榜梯度不存在。我们提出对评估的四项更改:三重评分、弃权(不执行)率报告、能力分层评估以及强制性校准指标。基准改革对于弥合该定理所识别的差距是必要的,但并非充分条件。

英文摘要

Three recent results describe what look like unrelated LLM reliability problems. Yin et al. (2026) show reasoning RL collapses tool-reliability representations. Suleymanov et al. (2026) show that under safety-constrained generation, large models rewrite flagged spans while small models truncate. Bastounis et al. (2024) prove any consistent-reasoning system without an implicit "I don't know" function must hallucinate infinitely often on broad problem classes. We argue these findings converge on a single intervention: calibrated abstention is what each independently identifies as the missing capability, even though the unavailability they document, a capability gap, a policy gap, and a recursion-theoretic gap, has a different source in each case. Honesty post-training has narrowed the gap in deployed models, but principled closure of the class Bastounis identifies requires a calibrated abstention function whose training signal at the leaderboard level is absent: dominant benchmarks assign zero reward to decline, so the leaderboard gradient that would select for the function does not exist. We propose four changes to evaluation: triple-scoring, abstention-rate reporting, capability-stratified evaluation, and mandatory calibration metrics. Benchmark reform is necessary, not sufficient, for closing the gap the theorem identifies.

发表机构

  • Amazon.com LLC(亚马逊公司)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑