arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.31245cs.CL

RupeeBias:审计大型语言模型在印度经济指导中的人口统计偏见

RupeeBias: Auditing Demographic Bias in Indian Economic Guidance from Large Language Models

Pavithra P M Nair, Bhavik Talaviya, Shourya Bhushan, Rahul Pankajakshan, Seema Guruvadoo, Avinash Agarwal, Gilad Gressel, Krishnashree Achuthan

首次发表
浏览论文内容

中文总结 AI 辅助

RupeeBias基准通过39,150个反事实提示审计LLM在印度经济指导中的偏见,覆盖六个轴线87个标识符,发现九个模型输出平均差异20.2%。

中文摘要 AI 辅助

个人在广泛的经济任务中求助于大型语言模型(LLM),从比较贷款选项、规划储蓄,到决定要求多少加薪或为自己的服务收取多少费用。众所周知,LLM会再现社会偏见,而有偏见的经济指导可能会影响用户对自己价值的认知、他们的要求以及他们最终接受的条件。这一风险在印度尤为突出,因为印度的经济结果受到诸如种姓和城乡位置等人口统计类别的影响。然而,现有的LLM偏见基准主要围绕西方人口统计类别设计,因此遗漏了印度背景下经济差距的关键轴线。我们引入了RupeeBias,一个用于在印度经济环境中审计LLM生成的经济指导中人口统计偏见的基准。RupeeBias包含39,150个提示,涵盖四个用例:薪资估算、薪资增幅估算、还价建议和服务定价建议。该基准采用单属性反事实设计,在改变一个人口统计标识符的同时,保持用户资质、经验或服务提供的描述不变。RupeeBias覆盖了六个轴线上的87个印度特定人口统计标识符:种姓、宗教、地区身份、性别、残疾和城乡位置,所有提示均以英语和印地英语(Hinglish)构建。我们在RupeeBias上评估了九个LLM,并发现所有六个轴线上都存在系统性的人口统计差异。对于仅在人口统计标识符上有所不同的、其他方面相同的提示,LLM生成的经济输出平均差异为20.2%。我们公开发布RupeeBias,以支持未来在印度特定人口统计和经济背景下对LLM生成的经济指导中人口统计偏见的研究。

英文摘要

Individuals turn to large language models (LLMs) for guidance across a wide range of economic tasks, from comparing loan options and planning savings to deciding what raise to ask for or how much to charge for their services. LLMs are known to reproduce social biases, and biased economic guidance may influence what users believe they are worth, what they ask for, and what they ultimately accept. This risk is especially salient in India, where economic outcomes are shaped by demographic categories such as caste and urban-rural location. Existing LLM bias benchmarks, however, are largely designed around Western demographic categories and therefore miss key axes of economic disparity in the Indian context. We introduce RupeeBias, a benchmark for auditing demographic bias in LLM-generated economic guidance across Indian economic settings. RupeeBias consists of 39,150 prompts spanning four use cases: salary estimation, salary increment estimation, counter-offer recommendation, and service pricing recommendation. The benchmark follows a single-attribute counterfactual design, holding the description of the user's qualifications, experience, or service offering fixed while varying one demographic identifier at a time. RupeeBias covers 87 India-specific demographic identifiers across six axes: caste, religion, regional identity, gender, disability, and urban-rural location, with all prompts constructed in both English and Hinglish. We evaluate nine LLMs on RupeeBias and find systematic demographic disparities across all six axes. For otherwise identical prompts that differ only in demographic identifier, LLM-generated economic outputs differ by 20.2% on average. We publicly release RupeeBias to support future research on demographic bias in LLM-generated economic guidance across India-specific demographic and economic contexts.

发表机构

  • Center for Cybersecurity Systems & Networks, Amrita Vishwa Vidyapeetham(阿姆里塔大学网络安全系统与网络中心)
  • Unique Identification Authority of India(印度唯一身份识别管理局)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑