发表机构
University of Luxembourg(卢森堡大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出道德熵框架,用贝叶斯后验分解标注分歧为偶然与认知不确定性,并审计共识规则,发现任意标注者规则产生约30%假阳性,多数规则遗漏63-83%真阳性。
AI 中文摘要
计算伦理学中的大多数工作将标注者在道德内容上的分歧视为应通过投票消除的噪声,一旦单个标注者标记某个项目,就将其合并为多数投票或更宽松的任意标注者规则。我们认为这种不确定性应当被建模并从中学习。我们引入了道德熵(Moral Entropy),这是一个贝叶斯框架,它保留关于真实标签的完整后验分布,并将其熵分解为偶然不确定性(关于道德内容的不可约分歧)和认知不确定性(源于标注不足或含噪),并允许任何启发式共识规则通过熵方法(如交叉熵/KL散度、布里尔分数和期望校准误差)对照校准后的真实标签进行审计。在三个语料库和十五个话语领域中,对照该后验审计标准聚合规则揭示了当前任何流程都未报告的偏差:任意标注者规则在大约30%的项目上与校准后验不一致——汇总来看,几乎全部是假阳性,尽管在基础层面错误发生反转(在MFTC上平均假阳性率/假阴性率为19.9%/38.9%)——而更严格的多数规则和两票规则遗漏了63-83%的真阳性。
英文摘要
Most work in computational ethics treats annotator disagreement on moral content as noise to be voted away, collapsed into majority vote or the more permissive any-annotator rule the moment a single annotator flags an item. We argue this uncertainty should instead be modeled and learned from. We introduce Moral Entropy, a Bayesian framework that keeps a full posterior over the true label and decomposes its entropy into aleatoric uncertainty (irreducible disagreement about the moral content) and epistemic uncertainty (from insufficient or noisy annotation) -- and lets any heuristic consensus rule be audited against a calibrated ground truth via entropy methods such as cross-entropy/KL, Brier score, and expected calibration error. Across three corpora and fifteen discourse domains, auditing the standard aggregation rules against this posterior reveals bias that no current pipeline reports: the any-annotator rule disagrees with the calibrated posterior on roughly 30% of items -- pooled, almost entirely false positives, though the errors invert at the foundation level (19.9%/38.9% mean FPR/FNR on MFTC) -- while the stricter majority and two-vote rules miss 63-83% of true positives.
Commentsaccepted to UncertaiNLP @ EMNLP 2026