arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

你的模型是在思考还是停滞不前?PUMA:通过相位动量对齐诊断推理病理学

UPAIR: Diagnosing Reasoning States via Uncertainty-Progress Alignment for Selective Intervention

Cheng Yan, Zhijun Fan, Guangyang Ye, Fan Xu, Xiang Xia, Yawei Wang, Wuyang Zhang

arXiv 2607.17188首次发表:更新:

发表机构

University of Science and Technology of China(中国科学技术大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究大型推理模型测试时的“过度思考”问题,提出相位动量对齐假设并构建认知能量模型,引入PUMA框架,通过分层诊断架构区分主动探索与被动停滞,在多基准实验中优于现有基线,实现更好的准确性-效率权衡和跨域泛化。

AI 中文摘要

测试时缩放使大型推理模型能够通过广泛的思维链来处理复杂任务。然而,这常常引发“过度思考”悖论,即冗余推理增加计算开销却无法保证准确性。现有测试时效率优化方法主要分两类:信息论方法易出现“欺骗性收敛”,潜在表征分析往往是事后的,缺乏对动态推理的实时敏感性。为弥补这一差距,我们提出相位动量对齐假设,理论上构建认知能量模型。接着引入PUMA,一个无需训练的框架,通过分层诊断架构有效区分主动探索与被动停滞,通过自适应截断或纠正措施实现精确干预。大量实验表明PUMA在不同基准上始终优于现有基线,实现了卓越的准确性-效率权衡和强大的跨域泛化。

英文摘要

While test-time scaling improves the problem-solving ability of large reasoning models (LRMs) through additional inference-time computation, it can also exacerbate overthinking and underthinking, which we formulate as reasoning state--action mismatch. Resolving this mismatch requires reliable reasoning state diagnosis, yet single-signal monitors provide ambiguous evidence, while steering-based controllers often rely on outcome-labeled supervision or model-specific calibration. We introduce the Uncertainty--Progress Alignment Hypothesis, which posits that the relative transition timing of proxy answer uncertainty and latent reasoning progress distinguishes healthy, stagnant, and ready states that warrant different subsequent actions. Building on this insight, we propose UPAIR, a training-free framework that couples lightweight uncertainty monitoring with event-triggered joint diagnosis and maps the resulting state to native continuation, selective strategy switching, or verification-guided stopping. Across three LRMs and five cross-domain benchmarks, the stagnation diagnosis detects 64.3% of natural errors while flagging only 5.4% of correct samples, revealing a dynamic reasoning regularity shared across models and tasks. End to end, UPAIR improves accuracy by up to 16.67 percentage points and reduces generated tokens by up to 29.64%, demonstrating the effectiveness of its integrated diagnosis and intervention, while online diagnosis costs less than 1% of natural-generation time.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑