arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.02089cs.LGcs.AI

推理摘要能透露多少信息?大语言模型的可观测性阶梯

How Much Does a Reasoning Summary Reveal? An Observability Ladder for Large Language Models

Andres Algaba, Francesca Carlon, Lynn Delcon, Marthe Ballon, Bert Verbruggen, Vincent Ginis

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出大语言模型的可观测性阶梯,通过实验发现有提示时摘要对正确性监测的帮助远小于完整轨迹,可监测性由显示内容和读者共同决定。

中文摘要 AI 辅助

大语言模型常向用户展示最终响应和简短推理摘要,而完整推理轨迹保持隐藏。本文提出一种可观测性阶梯,该阶梯固定每次已完成的运行,仅改变读者用于判断答案是否正确的检查内容,包括响应、模型从轨迹生成的自摘要、轨迹本身以及内部信号,每种内容分别在有提示和无提示的情况下进行测试。在三个基准和五个开放权重的Qwen3及gpt-oss模型上,我们针对每个访问级别训练匹配的线性正确性预测器。无提示时,摘要承载了轨迹的大部分排名信号(平均AUROC为0.774,而轨迹为0.813),且相比单独的响应提升了0.156;当提示可见时,摘要的增益降至0.019,而轨迹仍提升0.041。即使长度相同,轨迹的最后几个词预测正确性的效果与摘要相当或略好,且承载更密集、更具区分度的不确定性和自我修正线索。在同时存在正确和错误运行的MMLU-Pro问题上,线性摘要读者接近随机水平,轨迹读者仅保留适度信号,无论有无提示(无提示时AUROC为0.503-0.545,而轨迹为0.544-0.590)。无提示时,GPT-5-mini读者从gpt-oss-20b的摘要和轨迹中恢复了更多信号,即便如此,轨迹仍保持0.034的小幅优势。线性读者的轨迹信号大部分与长度相关。在用户已持有提示的常见情况下,摘要对监测正确性的帮助小于完整轨迹。因此,可监测性是显示内容与读者的共同属性,任何可监测性声明(包括忠实性声明)都应明确指定这两者。

英文摘要

Large language models often show users a final response and a short reasoning summary while the full reasoning trace stays hidden. We introduce an observability ladder that holds each completed run fixed and varies only what a reader inspects to judge whether the answer is correct: the response, a self-summary the model writes from the trace, the trace itself, and internal signals, each with and without the prompt. Across three benchmarks and five open-weight Qwen3 and gpt-oss models, we train matched linear correctness predictors on each access level. Without the prompt, summaries carry most of the trace's ranking signal (mean AUROC 0.774 versus 0.813) and add +0.156 over the response alone. With the prompt visible, the summary's gain collapses to +0.019, while the trace still adds +0.041. Even at equal length, the trace's last words predict correctness as well as summaries, or slightly better, and carry denser and more discriminative uncertainty and self-correction cues. On MMLU-Pro questions with both correct and incorrect runs, linear summary readers are near chance and trace readers retain only modest signal, both with and without the prompt (prompt-withheld AUROC 0.503-0.545 versus 0.544-0.590). With the prompt withheld, a GPT-5-mini reader recovers substantially more signal from both summaries and traces on gpt-oss-20b, and even then the trace keeps a small +0.034 advantage. Much of the linear readers' trace signal is associated with length. In the common case where users already hold the prompt, summaries are less helpful than the full trace for monitoring correctness. Monitorability is thus a joint property of the display and the reader, so any monitorability claim, including for faithfulness, should specify both.

补充信息

↑