arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

《“第10位陪审员”:面向官僚偏见检测的开放集立场筛选》

The "10th Juror": Open-Set Standpoint Screening for Bureaucratic Bias Detection

Yuchen Miao, Zijun Wang, Chang Han, Yurui Shi, Mingtai Zhang, Siyang Xu

arXiv 2610.11136首次发表:更新:

发表机构

Sydney Smart Technology College, Northeastern University; Taiyuan University of Technology; School of Resources, Environment and Materials, Guangxi University(东北大学悉尼智能科技学院; 太原理工大学; 广西大学资源环境与材料学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对荷兰政府文档的官僚偏见检测,提出MARS-Gov多智能体框架,引入动态“第10位陪审员”解决现有方法的三类缺陷,在DGDB数据集上取得最优性能。

AI 中文摘要

预设偏见的边界本身就是一种偏见。我们研究针对荷兰政府文档的闭环偏见治理,其中系统必须检测偏见性语言、将决策建立在法律和语境证据基础上、在需要干预时重写有问题的句子,并验证重写措施在不扭曲原意的前提下减轻了伤害。现有方法面临三项挑战:(i)判别式分类器捕捉表面规律但缺乏规范性依据;(ii)零样本大语言模型(LLMs)常采用通用视角,过度标记模糊的行政语言;(iii)固定分类法继承了闭世界假设,遗漏新兴的本地目标。我们提出MARS-Gov,这是一个感知立场的多智能体框架,结合了法律检索、开放集目标筛选、专门陪审员、保守路由和重写验证。当筛选发现未覆盖的群体时,MARS-Gov会实例化一个动态的“第10位陪审员”,在固定小组之外进行审议。在DGDB数据集上,MARS-Gov达到了新的最优性能(SOTA),F1值为0.880,比最强的零样本LLM检测器高出20.2个百分点(相对提升29.8%),比最佳的有监督荷兰编码器高出6.8个百分点,同时将不必要的干预降至2.5%。留一类别(LOCO)评估在保留类别上实现了85.1%的Correct@1和93.8%的Correct@3。

英文摘要

Presupposing the boundaries of bias is itself a form of bias. We study closed-loop bias governance for Dutch government documents, where a system must detect biased language, ground decisions in legal and contextual evidence, rewrite problematic sentences when intervention is warranted, and verify that the rewrite mitigates harm without distorting meaning. Existing methods face three challenges: (i) discriminative classifiers capture surface regularities but lack normative grounding; (ii) zero-shot LLMs often adopt generic viewpoints and over-flag ambiguous administrative language; and (iii) fixed taxonomies inherit the Closed-World Assumption, missing emerging local targets. We propose MARS-Gov, a standpoint-aware multi-agent framework that combines legal retrieval, open-set target screening, specialized jurors, conservative routing, and rewrite verification. When screening finds an uncovered group, MARS-Gov instantiates a dynamic "10th juror" to deliberate outside the fixed panel. On DGDB, MARS-Gov sets a new SOTA with 0.880 F1, outperforming the strongest zero-shot LLM detector by 20.2 points (29.8% relative) and the best supervised Dutch encoder by 6.8 points, while reducing unnecessary interventions to 2.5%. Leave-One-Category-Out (LOCO) evaluation recovers held-out categories with 85.1% Correct@1 and 93.8% Correct@3.

Comments19 pages, 6 figures. Accepted at EMNLP 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑