arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.01369cs.CL

你的回答有多正确?一个用于开放域问答评估的语义正确性框架

How Correct Is Your Answer? A Semantic Correctness Framework for Open QA Evaluation

Elitsa Yotkova, Violeta Kastreva, Petar Velkov, Hristo Boyanov, Dimitar Dimitrov, Ivan Koychev, Preslav Nakov

首次发表
浏览论文内容

中文总结 AI 辅助

针对开放域问答评估的瓶颈,该研究提出语义正确性分类法、相关基准数据集及CAP指标,CAP在单调性协议测试中表现优于现有基线。

中文摘要 AI 辅助

可靠的开放域问答评估仍是衡量现代大型语言模型(LLM)回答正确性的瓶颈。与多项选择任务不同,自由形式的答案可能有多种正确的表面形式,且可能存在不同性质的错误,包括不完整、矛盾、过度生成以及支持错误前提。现有的基于判断和相似度的指标往往忽略这些差异。我们通过三项可复用的贡献解决这一问题:首先,我们引入语义正确性分类法,将开放域答案分为8个有序类别,区分冗长但正确的答案与受幻觉内容污染的答案;其次,我们发布CAP-Correctness(一个包含8800个示例的基准,涵盖广泛使用的问答数据集)和CAP-Statements(一个包含11000个示例的数据集,用于将问答对转换为陈述性语句,以用于自然语言推理(NLI)训练和基于语句的评估);最后,我们引入CAP(上下文感知精确率),一种基于参考的指标,使用双向NLI对问题条件化的语句进行评分。在测试指标是否遵循分类法预期排序的单调性协议下,CAP的表现优于已有的基线方法。

英文摘要

Reliable evaluation of open-ended question answering remains a bottleneck for measuring answer correctness of modern LLMs. Unlike multiple-choice tasks, free-form answers may be correct in many surface forms and may fail in qualitatively different ways, including incompleteness, contradiction, overgeneration, and endorsement of false premises. Existing judgment-based and similarity-based metrics often collapse these distinctions. We address this gap with three reusable contributions. First, we introduce a semantic correctness taxonomy that assigns open-ended answers to eight ordered classes, separating verbose-but-correct answers from those contaminated by hallucinated content. Second, we release CAP-Correctness, an 8.8k-example benchmark spanning widely used QA datasets, and CAP-Statements, an 11k-example dataset for converting question-answer pairs into declarative statements for natural language inference (NLI) training and statement-based evaluation. Third, we introduce CAP (Context-Aware Precision), a reference-based metric that scores question-conditioned statements using bidirectional NLI. Under a monotonicity protocol testing whether metrics respect the taxonomy's intended ordering, CAP outperforms established baselines.

发表机构

  • Sofia University “St. Kliment Ohridski”(索非亚大学“圣克莱门特奥赫里德斯基”)
  • ETH Zürich(苏黎世联邦理工学院)
  • University of Zurich(苏黎世大学)
  • Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)

机构由 AI 辅助整理,请以论文原文为准。

↑