发表机构
ETH Zurich; UCLA(苏黎世联邦理工学院; 加州大学洛杉矶分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出CFG方法,利用因果推断从LLM生成中消除歧视性因果效应,同时保留合理路径,并在多个数据集上验证其有效性。
AI 中文摘要
大型语言模型(LLMs)越来越多地被用于生成、完善和转换信息,在这些场景中,其输出可能影响重大决策,从而引发对其对人口统计差异影响的担忧。在此背景下,因果推断为评估公平性提供了原则性基础,因为它将观察到的差异归因于产生这些差异的机制,而纯统计方法即使拥有无限数据也无法做到这一点。在LLM生成中,查询可能请求多个因果相关的变量,每个变量既是感兴趣的结果,又是其他输出的可能原因,并且提示中提供的信息不必遵循拓扑或时间顺序。这要求有方法能够分析并选择性地从这种灵活的生成过程中消除差异。在本文中,我们引入了基于大语言模型的因果公平生成(CFG,简称)。CFG提取相关概念,将生成基于参考人群和因果图,并消除用户选择的因果效应。CFG还允许保留那些被认为对任务效用合理的路径,这在法律文献中被称为“业务必要性”。此外,在适当的因果假设下,我们为在调整后的人群模型中消除所有歧视性因果效应提供了正式保证。我们使用四个LLM在三个基于人口数据的真实世界场景以及一个具有已知因果真实性的合成数据集上评估了CFG。
英文摘要
Large language models (LLMs) are increasingly used to generate, complete, and transform information in settings where their outputs can shape consequential decisions, raising concerns about their impact on demographic disparities. In this context, causal inference provides a principled basis for assessing fairness, because it attributes observed disparities to the mechanisms that generated them, which a purely statistical approach cannot do even with infinite data. In LLM generation, a query may request several causally related variables, each of which is both an outcome of interest and a possible cause of other outputs, and the information supplied in the prompt need not follow a topological or a temporal order. This calls for methods that can analyze and selectively remove disparities from such a flexible generation process. In this paper we introduce Causally Fair Generation with LLMs (CFG, for short). CFG extracts relevant concepts, grounds generation in a reference population and causal diagram, and removes user-selected causal effects. CFG also allows pathways deemed justifiable for the task's utility to be retained, which is known in legal literature as business necessity. Further, we provide formal guarantees for our method when eliminating all discriminatory causal effects in the adapted population model, under appropriate causal assumptions. We evaluate CFG with four LLMs in three real-world settings based on population data and on a synthetic dataset with a known causal ground truth.