arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.19678cs.CLcs.AIcs.LG

开放式问答中推理的无参考评估

Reference-Free Evaluation of Reasoning in Open-Ended Question Answering

发表机构伦敦大学玛丽皇后学院 · 天普大学
查看机构详情
  • Queen Mary University of London(伦敦大学玛丽皇后学院)
  • Temple University(天普大学)

机构由 AI 辅助整理,请以论文原文为准。

Guneet Singh Kohli, Yuxiang Zhou, Michael Sejr Schlichtkrull, Gregory E Dean, Maria Liakata

首次发表
浏览论文内容

中文总结 AI 辅助

针对高风险领域人工智能生成答案难验证问题,提出基于推理的无参考框架审核大语言模型输出,通过分解推理轨迹、标记关系并组织成超图等方法评估,在数学和医学推理设置下比直接以大语言模型为基线更可靠,强调问答评估应考虑推理关系组合。

中文摘要 AI 辅助

在高风险领域中,人工智能生成的答案往往流畅但难以验证,尤其是当它们包含多步推理而非单一最终答案时。我们提出了一个基于推理的无参考框架来审核大语言模型生成的输出。该方法将生成的推理轨迹分解为片段,使用自然语言推理标记局部前提-目标关系,并将这些关系组织成超图。然后,确定性反向与或搜索分配片段级审核标签,以表明每个片段在生成的响应中的基础方式。我们在两种设置下评估该框架:使用Hard2Verify进行演绎数学推理,以及使用UroReason进行开放式医学推理,UroReason是一个来自真实临床案例的新的医生注释的大语言模型推理轨迹基准。在这些设置中,我们的自然语言推理超图审核提供了比直接以大语言模型作为评判基线更可靠的无参考评估信号。在临床环境中,最先进的大语言模型评判往往无法识别有问题的推理片段,过度接受流畅但基础薄弱的响应。我们的结果表明,问答评估应考虑推理关系如何在推理轨迹中组合,而不是仅依赖最终答案或大语言模型作为验证者。UroReason将通过API提供,我们的代码将作为开源发布。

英文摘要

AI-generated answers in high-stakes domains are often fluent but difficult to verify, especially when they contain multi-step reasoning rather than a single final answer. We propose a reasoning-based, reference-free framework for auditing LLM-generated outputs. The method decomposes a generated reasoning trace into segments, labels local premise-target relations using Natural Language Inference (NLI), and organizes these relations into a hypergraph. A deterministic backward AND-OR search then assigns segment-level audit labels that indicate how each segment is grounded within the generated response. We evaluate the framework in two settings: deductive mathematical reasoning with Hard2Verify, and open-ended medical reasoning with UroReason, a new physician-annotated benchmark of LLM reasoning traces from real clinical cases. Across these settings, our NLI-hypergraph audit provides a more reliable reference-free evaluation signal than direct LLM-as-judge baselines. In the clinical setting, state-of-the-art LLM judges often fail to identify problematic reasoning segments, over-accepting fluent but weakly grounded responses. Our results show that QA evaluation should account for how inferential relations compose across a reasoning trace, rather than relying only on final answers or LLMs as verifiers. UroReason will be made available through an API, and our code will be released as open source.

↑