通过模型路由与协作实现集体偏见缓解
Collective Bias Mitigation via Model Routing and Collaboration
查看机构详情
- Nanyang Technological University(南洋理工大学)
- National University of Singapore(新加坡国立大学)
- Harbin Institute of Technology(哈尔滨工业大学)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
针对单一模型自我去偏不足的问题,提出集体偏见缓解框架,通过细粒度行为学习与多模型知识共享及路由协作,显著降低偏见得分,实现更公平的LLM响应。
中文摘要 AI 辅助
大型语言模型(LLMs)越来越多地部署于公共卫生、金融和治理领域,这要求其兼具准确性和社会价值对齐。尽管近期取得了进展,LLMs 常常会延续或放大其训练数据中固有的偏见,对公平性构成挑战。虽然自我去偏鼓励 LLM 识别并纠正自身偏见,但依赖单一模型的内在知识可能不足以解决根深蒂固的刻板印象。为解决这一局限,我们引入了集体偏见缓解(CBM),这是一种通过学习细粒度模型行为并促进不同 LLM 之间知识共享来缓解偏见的框架。本工作是首个系统探索不同 LLM 的有效选择与组织以培养更公平 LLM 响应的研究。实验表明,CBM 显著优于独立基线(例如,在 top-7 设置中,Committee 将年龄偏见得分从 0.25 降至 0.10)。我们的 Debating 和 Committee 拓扑结构实现了显著的偏见减少,其中后者在缓解效果与推理成本之间取得平衡,凸显了 CBM 在实现更公平 LLM 方面的潜力。
英文摘要
Large language models (LLMs) are increasingly deployed in public health, finance, and governance, requiring both accuracy and societal value alignment. Despite recent advances, LLMs often perpetuate or amplify bias embedded in their training data, posing challenges to fairness. While self-debiasing encourages an LLM to identify and correct its own biases, relying on a single model's intrinsic knowledge may be insufficient to address deeply ingrained stereotypes. To address this limitation, we introduce Collective Bias Mitigation (CBM), a framework that alleviates bias by learning fine-grained model behavior and fostering knowledge sharing among diverse LLMs. This work is the first to systematically explore the effective selection and organization of distinct LLMs to cultivate fairer LLM responses. Experiments show CBM substantially outperforms standalone baselines (e.g., in the top-7 setting, Committee lowers the age bias score from 0.25 to 0.10). Our Debating and Committee topologies achieve substantial bias reduction, with the latter balancing mitigation effectiveness and inference cost, highlighting the potential of CBM for fairer LLMs.