arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.05088cs.AIcs.CLcs.CY

通过论证分析衡量AI问责制:模型推理能否经受审查?

Measuring AI Accountability Through Argumentation Analysis: Can Model Reasoning Withstand Scrutiny?

Daan R. Henselmans, Derck W. E. Prinzhorn, Arno Libert

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出基于Walton和Govier理论的四阶段辩证协议,通过9个前沿模型在200个高模糊性道德选择任务上的实验,衡量AI裁决辩护的结构质量,揭示了模型推理与辩护方案的差异及对齐评估的情境性需求。

中文摘要 AI 辅助

AI监督方法依赖真实值进行验证,但何种行为构成恰当的AI行为存在争议。这使得对大型语言模型(LLMs)道德推理的评估,以及基于辩论的监督,都在无形中避开了现实中的模糊性。我们研究了一种在这种模糊性下仍能发挥作用的替代标准:通过四阶段辩证协议衡量模型针对关键问题为其裁决所提供辩护的结构质量,该协议基于Walton的论证方案理论和Govier的论证说服力标准。该协议可适应不同的推理框架,超出了多项选择框架的限制,同时兼顾裁决前的推理及其事后辩护。在9个前沿模型和200个高模糊性MoralChoice项目(共6778个评判者评分单元,经二元失败判断的评判者间一致性达89.6%验证)中,各模型在所有维度上的辩护表现均高于标准最低要求。失败集中在理由和充分性方面,且与认知模糊性相关,而非论证长度。在所有模型和所有Govier维度上,推理的辩护都优于事后辩护。尽管基于价值的实践推理在两条路径中占主导,但模型在辩护中呈现的方案与其推理所用方案,在相当一部分困境(每个模型均≥20%)上存在差异。该协议能捕捉到完全站不住脚的辩护(自相矛盾、虚假前提),并揭示了在AI对齐中刻画弃权(不执行)作用的困难,表明需要更具情境性的评估。

英文摘要

AI oversight methods rely on ground truth for validation, but what constitutes appropriate AI behavior is contested. This leaves evaluation of moral reasoning in LLMs and debate-based oversight implicitly avoiding realistic ambiguity. We investigate an alternative standard designed to function despite such ambiguity: structural quality of the defence a model can mount for its verdicts in response to critical questions, measured through a four-phase dialectical protocol grounded in Walton's theory of argumentation schemes and Govier's criteria for argument cogency. The protocol is adaptive to different frames of reasoning, extends beyond multiple-choice framing, and treats both the reasoning that precedes a verdict and its post-hoc justification. Across nine frontier models and 200 high-ambiguity MoralChoice items -- $6,778$ judge-scored cells, validated against $89.6\%$ inter-judge agreement on the binary failure judgment -- models defend their reasoning well above the rubric minimum on every dimension. Failure mass concentrates on grounds and sufficiency, and correlates with epistemic hedging rather than argument length. Reasoning is better defended than post-hoc justification, on every model and every Govier dimension. The scheme a model presents in its justification differs from the one it reasoned with on a substantial share of dilemmas ($\geq 20\%$ per model), despite value-based practical reasoning dominating both tracks. The protocol catches strictly indefensible defences (self-contradiction, false premises), and it surfaces difficulties in characterizing the role of retraction in AI alignment, suggesting a need for more situated evaluations.

发表机构

  • Aithos Research Foundation(艾索斯研究基金会)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑