相同问题,不同答案?衡量与缓解提示特权以实现公平的AI访问
Same Question, Different Answer? Measuring and Mitigating Prompt Privilege for Equitable AI Access
- Fuqua School of Business Duke University(杜克大学福库商学院)
- Department of Engineering Carnegie Mellon University(卡内基梅隆大学工程学院)
- Department of Industrial Engineering and Management Sciences Northwestern University(西北大学工业工程与管理科学系)
- Department of Information and Decision Sciences University of Minnesota(明尼苏达大学信息与决策科学学院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
该研究提出统一框架,引入PES指标与PET智能体,在MedQA基准上验证了提示特权的存在,PET可消除该差异,提升AI访问公平性。
AI中文摘要:
大型语言模型(LLMs)正越来越多地被集成到医疗保健、教育、公共服务及日常决策中,无论用户的读写能力、沟通风格或提示工程专业水平如何,它们都应提供相当的帮助。然而,现有的关于提示鲁棒性的研究主要聚焦于对抗性攻击、提示注入和提示优化,却忽略了语义等价的请求是否仅因表述方式不同而得到不同的响应。我们将这一可访问性挑战称为“提示特权”:具备更强提示专业知识的用户,即便表达的是相同的核心意图,也会系统性地获得更好的模型性能。为解决该问题,我们提出了一个用于衡量和缓解LLM交互中可访问性差异的统一框架。我们引入了提示公平得分(Prompt Equity Score, PES),这是一种用于评估不同用户群体间性能一致性的量化指标;还提出了提示公平转换器(Prompt Equity Transformer, PET),这是一种基于LLM的智能体,可自动将用户请求转换为语义等价、面向可访问性的提示,同时保留其意图。PET将提示优化从用户转移到AI系统,充当用户与基础模型之间的智能可访问性层。在MedQA基准上的实验显示出可测量的提示特权,低读写能力用户群体与专家提示用户群体之间存在具有统计显著性的性能差异。应用PET可消除这些差异,同时保留语义保真度,表明面向可访问性的提示规范化可改善公平的AI访问。通过将提示特权作为AI可访问性的新维度引入,并将PET作为实用解决方案,本研究推进了以系统为中心的可访问性,并为构建更公平、可信和包容的AI系统奠定了基础。
英文摘要:
Large language models (LLMs) are increasingly integrated into healthcare, education, public services, and everyday decision making. They should provide comparable assistance regardless of a user's literacy, communication style, or prompt-engineering expertise. However, existing research on prompt robustness primarily focuses on adversarial attacks, prompt injection, and prompt optimization, while overlooking whether semantically equivalent requests receive different responses simply because they are phrased differently. We refer to this accessibility challenge as "Prompt Privilege": users with greater prompting expertise systematically obtain better model performance despite expressing the same underlying intent. To address this problem, we present a unified framework for measuring and mitigating accessibility disparities in LLM interactions. We introduce Prompt Equity Score (PES), a quantitative metric for evaluating performance consistency across user populations, and Prompt Equity Transformer (PET), an LLM-based agent that automatically transforms user requests into semantically equivalent, accessibility-oriented prompts while preserving their intent. PET shifts prompt optimization from the user to the AI system, functioning as an intelligent accessibility layer between users and foundation models. Experiments on the MedQA benchmark demonstrate measurable prompt privilege, with statistically significant performance disparities between low-literacy and expert-prompting cohorts. Applying PET eliminates these disparities while preserving semantic fidelity, demonstrating that accessibility-oriented prompt normalization can improve equitable AI access. By introducing prompt privilege as a new dimension of AI accessibility and PET as a practical solution, this work advances system-centered accessibility and provides a foundation for more fair, trustworthy, and inclusive AI systems.