arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

省略行动:哲学分歧下衡量框架不变的省略偏差

OMIT the Action: Measuring Framing-Invariant Omission Bias under Philosophical Disagreement

Sihyeon Lee, Jihun Song, Chanwoo Kim, Jiwoo Kum, Chanjun Park

arXiv 2610.07847首次发表:更新:

发表机构

Soongsil University(崇实大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对LLM道德推理中的省略偏差,构建OMIT基准(218个配对框架场景),评估发现偏差普遍且与模型规模负相关,并验证了推理时干预可减少偏差、增强框架一致性。

AI 中文摘要

随着大型语言模型(LLM)越来越多地参与道德推理,省略偏差(即倾向于不行动,即使等效的框架反转了实质性结果)对决策的公正性构成了显著风险。然而,省略偏差在LLM评估中仍未得到充分探索,现有的少数研究规模有限,且主要集中在功利主义与义务论冲突上。为弥补这一空白,我们引入了OMIT基准,该基准包含10种冲突类型下的218个配对框架场景,通过利用基于LLM的五视角哲学人格小组(功利主义、义务论、美德伦理、关怀伦理和契约主义)中的分歧模式构建而成。评估八个LLM后,我们发现省略偏差普遍存在,但在同一模型家族内与模型规模呈负相关。我们进一步评估了四种推理时干预措施,发现鼓励模型在给出是/否答案前考虑道德原则的干预措施能减少省略偏差并增加框架一致的反应,尽管较低的省略偏差率也可能伴随向行动偏差反应的转变。最终,这项工作不仅贡献了OMIT基准,还提供了一种方法,利用多样化的哲学分歧信号来评估框架敏感的“不行动”偏好,以及LLM在复杂道德冲突下缓解尝试的分布效应。

英文摘要

As LLMs increasingly assist in moral reasoning, omission bias, the tendency to prefer inaction even when equivalent framings reverse substantive outcomes, poses a significant risk of skewed decision-making. Yet omission bias remains underexplored in LLM evaluation, with the few existing studies limited in scale and focused largely on utilitarian-deontological conflicts. To address this gap, we introduce OMIT, a benchmark consisting of 218 paired-frame scenarios across 10 conflict types, constructed by leveraging disagreement patterns from an LLM-based, five-perspective philosophical persona panel (utilitarianism, deontology, virtue ethics, care ethics, and contractualism). Evaluating eight LLMs, we find that omission bias is pervasive but inversely correlates with model size within families. We further evaluate four inference-time interventions and find that interventions encouraging models to consider moral principles before committing to a yes/no answer reduce omission bias and increase frame-consistent responses, although lower omission bias rates can also coincide with shifts toward action-biased responses. Ultimately, this work contributes not only the OMIT benchmark, but also a methodology for using diverse philosophical disagreement signals to evaluate framing-sensitive inaction preferences and the distributional effects of mitigation attempts in LLMs under complex moral conflicts.

CommentsAccepted to AACL-IJCNLP 2026 Findings

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑