arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.11994cs.AIcs.CL

面向高效测试时推理的声明级可靠性评估

Claim-Level Reliability Assessment for Efficient Test-Time Reasoning

Sen Xu, Wei Wang, Shixi Liu, Jixin Min, Yingwei Dai, Zhibin Yin, Yirong Chen, Junlin Zhang

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出无需训练的CLR框架,将测试时计算从解决方案采样重分配至声明级语义证伪,在匹配预算下于四个LLM及四个推理基准上提升了推理性能,减少了标记使用量。

中文摘要 AI 辅助

我们提出声明级证伪作为测试时扩展的原则,并通过声明级可靠性评估(Claim-Level Reliability Assessment, CLR)将其实例化,这是一个无需训练的框架,可将测试时的计算从额外的解决方案采样重新分配到针对性的验证。由于完整推理轨迹的评估常因常规标记的信号稀释而掩盖决定性错误,CLR将每条推理轨迹浓缩为一组紧凑的决策关键声明,从而分离其逻辑锚点。此外,考虑到在固定模型能力下生成完全正确的解决方案固有难度,CLR将重点转向语义证伪。该方法利用了解决方案构建与声明反驳之间的根本不对称性:构建有效解决方案需要完美的推理路径,而反驳错误声明仅需识别单个决定性缺陷。这种对负面证据的针对性搜索系统性压缩了高置信度错误轨迹的生存空间,通过非线性可靠性评分有效抑制了错误共识。在匹配预算下,针对四个大型语言模型(LLM)和四个推理基准,CLR总体上在pass@1和自一致性指标上实现了提升。例如,在GPT-OSS-20B/CMIMC25上,CLR使pass@1超出27.15个百分点,并将自一致性准确率从77.50%提升至82.19%,同时减少了37.0%的标记使用量。

英文摘要

We propose claim-level falsification as a principle for test-time scaling and instantiate it through Claim-Level Reliability Assessment (CLR), a training-free framework that reallocates test-time compute from additional solution sampling to targeted verification. Since whole-trace evaluation often obscures decisive errors due to signal dilution from routine tokens, CLR condenses each reasoning trace into a compact set of decision-critical claims, thereby isolating its logical anchors. Furthermore, recognizing the inherent difficulty of generating entirely correct solutions under fixed model capabilities, CLR shifts the focus to semantic falsification. This approach exploits a fundamental asymmetry between solution construction and claim refutation. Constructing a valid solution requires a flawless reasoning path, whereas refuting an incorrect claim requires identifying only a single decisive flaw. This targeted search for negative evidence systematically compresses the survival space of high-confidence incorrect traces, effectively suppressing erroneous consensus via nonlinear reliability scoring. Across four LLMs and four reasoning benchmarks under matched budgets, CLR generally improves upon pass@1 and self-consistency. On GPT-OSS-20B/CMIMC25, for instance, CLR exceeds pass@1 by 27.15 percentage-points and raises self-consistency accuracy from 77.50\% to 82.19\% with 37.0\% fewer tokens.

发表机构

  • Sina Weibo Inc.(新浪微博公司)

机构由 AI 辅助整理,请以论文原文为准。

↑