arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

鲁棒去中心化公平性审计

Robust Decentralized Fairness Auditing

Sayan Biswas, Jade Garcia Bourrée, Anne-Marie Kermarrec, Palak, Martijn de Vos

arXiv 2610.10199首次发表:更新:

发表机构

EPFL(洛桑联邦理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对多审计者协作审计LLM公平性时的公平洗白攻击,提出Auditopus去中心化审计方法,通过统计一致性降权防御,显著降低审计错误,即使49%审计者对抗也不放过不公平模型。

AI 中文摘要

新兴法规要求对大型语言模型(LLM)进行审计,以检查其是否符合监管标准,尤其是公平性。此类黑盒审计通常假设存在一个能够访问大量具有代表性查询集的单一审计者。在实践中,单一审计者可能难以获得这样的查询集,但多个审计者可以通过使用各自的查询集协作审计LLM,共同覆盖相关的 demographic 群体。然而,依赖多个审计者会引发一个根本性的信任问题,因为他们可能代表LLM提供方行事,以制造公平的误导性表象,即“公平洗白”。我们提出了Auditopus,一种用于鲁棒去中心化公平性审计的新方法。在Auditopus中,审计在没有中央服务器的情况下按轮进行。在每一轮中,每个审计者向LLM发出固定数量的查询,并仅将其查询结果的累积统计向量发送给其他审计者,而不是以明文形式发送敏感的查询。然后通过聚合所有向量来估计被审计LLM的公平性。我们在理论上和实证上表明,即使网络中的单个对抗性审计者也可以通过伪造其发送的向量来操纵这种估计,使不公平的LLM看起来公平。为了应对这一威胁,Auditopus让每个诚实审计者在本地对任何其累积统计向量与先前向量在统计上不一致的审计者进行降权。我们实现了Auditopus,并在两个数据集上使用两个预训练的LLM,将其与鲁棒聚合基线进行比较。针对一个优化其发送向量以使LLM看起来公平的攻击者,Auditopus相对于无防御平均将审计错误降低高达78%,相对于鲁棒聚合基线至少降低62%。即使当49%的审计者是对抗性的,Auditopus也绝不允许非常不公平或中等不公平的LLM通过为公平。

英文摘要

Emerging legislation requires large language models (LLMs) to be audited for compliance with regulatory standards, particularly fairness. Such black-box audits typically assume a single auditor with access to a large, representative set of queries. In practice, it can be difficult for an auditor to obtain such a query set, but multiple auditors can together cover the relevant demographic groups by auditing the LLM collaboratively with their individual query sets. However, relying on multiple auditors raises a fundamental trust problem, as they may act on behalf of the LLM provider to portray a misleading appearance of fairness, i.e., fairwashing. We propose Auditopus, a novel approach for robust decentralized fairness auditing. In Auditopus, auditing proceeds in rounds without a central server. In each round, every auditor issues a fixed number of queries to the LLM, and sends only cumulative statistics vectors of its query results to other auditors instead of sensitive queries in clear. The fairness of the audited LLM is then estimated by aggregating all the vectors. We show theoretically and empirically that even a single adversarial auditor in the network can steer this estimate by fabricating the vectors it sends, making an unfair LLM appear fair. To address this threat, Auditopus has each honest auditor locally down-weight any auditor whose cumulative statistics vectors are statistically inconsistent with previous ones. We implement Auditopus and compare it to robust aggregation baselines on two datasets with two pre-trained LLMs. Against an attacker that optimizes the vectors it sends to make the LLM appear fair, Auditopus reduces audit error by up to 78% on average relative to no defense and at least 62% relative to the robust aggregation baselines. Even when 49% of the auditors are adversarial, Auditopus never lets a very unfair or moderately unfair LLM pass as fair.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑