arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

错误预测,正确答案:从崩溃的大语言模型序列得分中恢复证据

Wrong Prediction, Right Answer: Recovering Evidence from Collapsed LLM Sequence Scores

Qiyao Yan, Chenpeng Wang, Liangming Pan

arXiv 2608.31068首次发表:更新:

发表机构

Peking University; YiXin-AILab; Beijing Academy of Artificial Intelligence(北京大学; 艺心人工智能实验室; 北京人工智能研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究发现大语言模型推理失败多为表达缺陷而非逻辑缺失,提出无目标标签的加性校正方法,可恢复Qwen等模型的推理准确率,为基准评估提供更严谨解读。

AI 中文摘要

当大语言模型在推理任务中失败时,通常会被认为缺乏底层能力,但这混淆了真正的推理缺失与后期输出瓶颈。我们在多种推理基准中观察到一致的读出差距:即使原生序列评分因结构偏差完全崩溃,隐藏状态探针仍能成功解码正确答案。为测试特定实例的逻辑是否在崩溃中保留,我们引入了一种使用最小化、无目标标签的加性校正的诊断协议。仅用25个未标记示例拟合2个参数,就能为Qwen3.5模型恢复9至34个准确率点,且可成功迁移至OLMo-2-1B和Llama-3.1-8B。关键是,这些恢复的决策在简单词汇重叠无法解决的困难实例上依然存在,且显著优于保持计数的置换基线。我们的结果表明,许多表面上的零样本推理缺陷是表达失败掩盖了完整的内部逻辑,这要求对基准评估采取更严谨的解读。

英文摘要

When a large language model fails a reasoning task, it is often assumed to lack the underlying capability. However, this conflates a genuine absence of reasoning with a late-stage output bottleneck. We observe a consistent readout gap across diverse reasoning benchmarks: hidden-state probes successfully decode correct answers even when native sequence scoring completely collapses due to structural biases. To test whether instance-specific logic survives this collapse, we introduce a diagnostic protocol using a minimal, target-label-free additive correction. Fitting just two parameters on as few as 25 unlabeled examples recovers 9--34 accuracy points for Qwen3.5 models, transferring successfully to OLMo-2-1B and Llama-3.1-8B. Crucially, these recovered decisions persist on hard instances unresolved by simple lexical overlap and significantly exceed count-preserving permutation baselines. Our results show that many apparent zero-shot reasoning deficits are expression failures masking intact internal logic, urging a narrower interpretation of benchmark evaluations.

Comments20 pages, 4 figures, and 43 appendix tables. Research paper on language-model interpretability, reasoning evaluation, output-scoring bottlenecks, and label-free calibration. The paper evaluates controlled three-way logical reasoning tasks, ProofWriter, ANLI, and FOLIO using Qwen3.5, OLMo-2-1B, Llama-3.1-8B, and Pythia checkpoints

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑