arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.27499cs.AI

已标注评分的概念论证数据集

A dataset of rated conceptual arguments

Emery Cooper, Caspar Oesterheld, Linh Chi Nguyen, Alexander Kastner, Ethan Perez

首次发表
浏览论文内容

中文总结 AI 辅助

该研究构建了含多维度专家评分的概念论证数据集,提出两种评分函数并基准测试多模型,发现模型性能与通用能力排名相关。

中文摘要 AI 辅助

大型语言模型在数学、编程等有可验证答案的任务上已取得快速进步,但其在概念性问题上的推理能力却鲜为人知:这类问题没有现实可及的真实答案,也不存在广泛认可的解决方法,但可通过辩论论证取得进展,多数哲学问题、AI安全、决策理论及社会选择领域的核心问题均属此类。我们认为,虽这类问题的最终结论难以评估,但可更可靠地评估单个情境化论证。为此,我们引入一个包含951条论证性评论的数据集,这些评论针对442篇立场文本,涵盖AI安全、决策理论、伦理及政治等主题,由6位专家评分者从中心性、强度、正确性、清晰度等维度给出1458个评分。我们提出两种评分函数,并对一系列模型进行基准测试,结果显示模型性能与通用能力排名相符。

英文摘要

Large language models have improved rapidly on tasks with verifiable answers, such as mathematics and programming. Much less is known about their ability to reason about what we call conceptual questions: questions for which no ground truth is realistically accessible and no widely accepted resolution methodology exists, but on which progress can still be made by debating arguments. Most philosophical questions are of this kind, as are central components of questions in AI safety, decision theory, and social choice. Our approach is based on the view that while bottom-line conclusions on such questions are hard to evaluate, individual contextualized arguments can be evaluated far more reliably. We therefore introduce a dataset of 951 argumentative critiques of 442 position texts, spanning topics from AI safety and decision theory to ethics and politics, with 1,458 ratings by six expert raters along dimensions including centrality, strength, correctness, and clarity. We propose two scoring functions and benchmark a range of models. Performance tracks general capability rankings.

↑