发表机构
University of North Texas; University of Alaska Anchorage(北德克萨斯大学; 阿拉斯加大学安克雷奇分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出ArcticQA数据集与ArcticAbstain基准,评估LLM在北极科学多选题中的弃权(不执行)行为,发现答案缺失时弃权率平均提升5.05个百分点,强调需联合评估弃权频率与响应性。
AI 中文摘要
大型语言模型(LLMs)在面对科学多选题时,若没有选项有效,应弃权(不执行),但仅频繁弃权并不能证明模型对答案可用性的敏感性。我们引入了ArcticQA,一个包含194个问题、源自北极原始研究的数据集,并针对答案支持与干扰项矛盾进行了基于源证据的自动化检查。我们进一步开发了ArcticAbstain,一个配对基准,比较答案存在与答案缺失两种条件,后者将正确答案替换为干扰项,并在两种条件下均提供明确的弃权(不执行)选项。我们评估了Gemini、Claude和ChatGPT系列中的八个模型,在高推理努力下,每种条件进行三次试验,共记录了9,312个响应。答案存在条件下的弃权(不执行)率介于0.0%至63.0%之间,而替换正确答案后,弃权(不执行)率平均增加5.05个百分点。这些发现凸显了显著的基线差异,以及需要联合评估弃权(不执行)频率与响应性的必要性。数据集和基准可在该https URL获取。
英文摘要
Large language models (LLMs) should abstain from scientific multiple-choice questions when no option is valid, but frequent abstention alone does not demonstrate sensitivity to answer availability. We introduce ArcticQA, a dataset of 194 questions derived from primary Arctic research, with automated checks of answer support and distractor contradiction against source evidence. We further develop ArcticAbstain, a paired benchmark comparing answer-present and answer-absent conditions, with the correct answer replaced by a distractor in the latter and an explicit abstention option in both. We evaluate eight models from the Gemini, Claude, and ChatGPT families at high reasoning effort, with three trials per condition, yielding 9,312 recorded responses. Answer-present abstention rates range from 0.0% to 63.0%, whereas replacing the correct answer increases abstention by 5.05 percentage points on average. These findings highlight substantial baseline differences and the need to evaluate abstention frequency and responsiveness jointly. The dataset and benchmark are available at https://github.com/BenWilcox8/arctic-qa.