arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

检测大语言模型中沉默推理失败的无参考分数

A Reference-Free Score for Detecting Silent Reasoning Failures in Large Language Models

Vivek Shukla, Varun Shukla, Atul, Divya Mishra, Mehul Kumar Das

arXiv 2607.26102首次发表:更新:

发表机构

Allenhouse Institute of Technology(艾伦豪斯技术学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对大语言模型数学思维链评估的推理-答案一致性差距问题,提出无参考的RAFS分数,结合多维度指标诊断沉默推理与答案提取错误,相关研究在GSM8K、MATH数据集上开展验证。

AI 中文摘要

数学思维链(CoT)评估通常简化为最终答案是否与参考答案匹配,这将得出正确结论与生成有效推导混为一谈:无效的思维链可能偶然得出正确答案,而有效的计算之后可能出现转录错误,我们将这种不匹配称为推理-答案一致性差距。本框架论文提出推理-答案可信度分数(RAFS),这是一种无参考、实例级别的诊断指标,用于判断输出的数学轨迹是否局部可信、是否支持其答案,以及在重采样和针对性反事实干预下是否稳定。RAFS结合了步骤有效性、推理对答案的蕴含关系、反事实敏感性、答案一致性和条件推理稳定性,它评估的是转录层面的一致性,而非模型的内部计算,也不评估所测试数学场景之外的事实正确性。我们保留了一项预先注册、结果盲法的验证性研究,该研究针对GSM8K和MATH数据集,其假设、可接受性规则、校准和测试均在检查验证性结果前确定;还指定了一项单独的可行性试点,用于验证端到端执行情况并估计干预覆盖率,在该试点冻结后才报告数值试点结果,且仅在有轨迹级人工制品时才报告。我们将四种推理-答案结果形式化,论证非补偿性聚合器的合理性,实例化语义轨迹距离,量化计算与弃权(不执行)的权衡,并定义验证者独立性和功效分析。RAFS旨在作为数学答案准确性的补充,为沉默推理失败和答案提取错误提供可审计的警告信号。

英文摘要

Mathematical chain of thought (CoT) evaluation is commonly reduced to whether the final answer matches a reference. This conflates producing a correct conclusion with producing a valid derivation an invalid chain can accidentally reach the right answer, while a valid calculation can be followed by a transcription error. We call this mismatch the reasoning answer consistency gap. This framework paper introduces the Reasoning Answer Faithfulness Score (RAFS), a reference free, instance level diagnostic of whether an emitted mathematical trace is locally credible, supports its answer, and is stable under resampling and targeted counterfactual interventions. RAFS combines step validity, reasoning to answer entailment and counterfactual sensitivity, answer consensus, and conditional reasoning stability. It evaluates transcript level agreement, not a models private computation and not factual correctness outside the tested mathematical setting. We retain a preregistered, results blind confirmatory study on GSM8K and MATH, with hypotheses, admissibility rules, calibration, and tests fixed before confirmatory outcomes are inspected. A separate feasibility pilot is specified to verify end to end execution and estimate interven tion coverage before that freeze numerical pilot claims are re ported only when trace level artifacts are available. We formalize four reasoning answer outcomes, justify the non compensatory aggregator, instantiate semantic trace distance, quantify compute and abstention tradeoffs, and define verifier independence and power analyses. RAFS is intended to complement mathematical answer accuracy with an auditable warning signal for silent reasoning failures and answer extraction errors

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑