发表机构
University of Georgia; University of Southern California; Carnegie Mellon University; Michigan State University(佐治亚大学; 南加州大学; 卡内基梅隆大学; 密歇根州立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出首个联合评估LLM推荐器可观测输出与隐藏表示公平性的基准FairGap,发现二者存在根本张力,现有框架无法诊断。
AI 中文摘要
针对基于大语言模型(LLM)的推荐器的公平性审计大多聚焦于可观测输出,隐含假设稳定的推荐结果反映了稳定的内部处理过程。我们通过FairGap对该假设提出挑战,FairGap是首个在两个层面联合评估推荐公平性的基准:可观测输出偏移(OBS)与隐藏表示偏移(IBS),通过针对性别、年龄和种族的受控反事实身份探测进行测量。二者的关系通过表示-输出对齐(ROA)进行总结,带有象限诊断以识别用户级隐藏-输出不匹配。FairGap应用于三个领域的六个开源权重LLM家族,揭示了普遍存在的隐藏-输出解耦:ROA极少超过0.22,且有相当一部分用户群体尽管内部发生了显著偏移,仍表现出稳定的输出,这是仅输出审计因设计无法检测的模式。此外,可将IBS降低多达8倍的激活引导会同时恶化OBS,表明内部与输出级公平性之间存在现有框架无法诊断的根本张力。
英文摘要
Fairness audits for LLM-based recommenders have largely focused on observable outputs, implicitly assuming that stable recommendations reflect stable internal processing. We challenge this assumption with FairGap, the first benchmark to jointly evaluate recommendation fairness at two levels: observable output shift (OBS) and hidden representation shift (IBS), measured through controlled counterfactual identity probes across gender, age, and race. Their relationship is summarized via Representation-Output Alignment (ROA), with quadrant diagnostics for identifying user-level hidden-output mismatch. Applied to six open-weight LLM families across three domains, FairGap reveals pervasive hidden-output decoupling: ROA rarely exceeds 0.22, and a non-negligible user population shows stable outputs despite substantial internal shifts, a mode that output-only audits cannot detect by design. Further, activation steering that reduces IBS by up to 8x simultaneously worsens OBS, demonstrating a fundamental tension between internal and output-level fairness that existing frameworks are unequipped to diagnose.