Wrong Prediction, Right Answer: Recovering Evidence from Collapsed LLM Sequence Scores
错误预测,正确答案:从崩溃的大语言模型序列得分中恢复证据
机构 * Peking University(北京大学) ; YiXin-AILab(艺心人工智能实验室) ; Beijing Academy of Artificial Intelligence(北京人工智能研究院)
AI总结 该研究发现大语言模型推理失败多为表达缺陷而非逻辑缺失,提出无目标标签的加性校正方法,可恢复Qwen等模型的推理准确率,为基准评估提供更严谨解读。
Comments 20 pages, 4 figures, and 43 appendix tables. Research paper on language-model interpretability, reasoning evaluation, output-scoring bottlenecks, and label-free calibration. The paper evaluates controlled three-way logical reasoning tasks, ProofWriter, ANLI, and FOLIO using Qwen3.5, OLMo-2-1B, Llama-3.1-8B, and Pythia checkpoints