AISPA:面向大语言模型应用的以用户为中心的系统提示审计框架
AISPA: User-Centric System Prompt Auditing for Large Language Model Applications
- Stanford University(斯坦福大学)
- CMU(卡内基梅隆大学)
- UT Austin(德克萨斯大学奥斯汀分校)
- University of Toronto(多伦多大学)
- UCSB(加利福尼亚大学圣巴巴拉分校)
- WashU(华盛顿大学)
- OSU(俄亥俄州立大学)
- UCSC(加利福尼亚大学圣克鲁兹分校)
- Northwestern University(西北大学)
- UIUC(伊利诺伊大学厄巴纳-香槟分校)
- KAUST(阿卜杜拉国王科技大学)
- MIT(麻省理工学院)
- University of Oxford(牛津大学)
- Institute for Decentralized AI(去中心化人工智能研究所)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出以用户为中心的AISPA框架,审计88款商业AI产品的3249条系统提示指令,发现其设计差异大、保护指令范围浅、长度增长但仍存问题指令,凸显系统提示需更高透明度与监督。
AI中文摘要:
系统提示是开发者配置的用于控制AI应用中基础模型行为的指令,广泛应用于商业AI产品,但很少向公众或监管机构披露,这在AI系统的广泛部署中造成了严重的信任和问责缺口。本文提出人工智能系统提示保障(AISPA),这是一个以用户为中心的、用于系统审计AI系统中系统提示的框架。AISPA会检查系统提示的特定部分,并沿8个对用户重要的维度对其进行评估。随后,我们使用该框架审查了88款商业AI产品中系统提示的3249条指令,将每条指令归类为保护用户的或存在问题的。我们的审计得出四项核心发现:第一,不同产品和开发者的系统提示设计差异巨大,部分组织平均每个产品有超过60条保护指令,而另一些组织平均不到5条;第二,保护指令被广泛采用,但范围较浅:98.9%的产品至少包含一条保护指令,但仅有24%的产品覆盖了AISPA分类法的全部8个维度;第三,系统提示的长度稳步增长,且对用户的保护力度不断增强,这表明用户保护正成为商业提示设计中更受关注的问题;第四,尽管取得了上述进展,存在问题的指令仍然普遍存在:约40%的产品至少包含一条损害用户利益的指令,且保护指令与存在问题的指令经常共存于同一系统提示中。我们的研究结果强调,商业AI产品的系统提示需要更高的透明度、标准化和独立监督。
英文摘要:
System prompts are instructions configured by developers to govern the behaviors of foundation models in AI applications. They are used throughout commercial AI products, but are rarely disclosed to the public or regulators, creating a serious trust and accountability gap in the wide deployment of AI systems. In this paper, we introduce Artificial Intelligence System Prompt Assurance (AISPA), a user-centric framework for systematically auditing system prompts in AI systems. AISPA examines specific parts of a system prompt and evaluates them along eight dimensions that matter to users. We then use this framework to review 3,249 instructions from system prompts in 88 commercial AI products, classifying each instruction as either protective (of users) or problematic. Our audit surfaces four core findings. First, system prompt design varies substantially across products and developers, with some organizations averaging over 60 protective instructions per product while others average fewer than 5. Second, protective instructions are widely adopted but shallow in scope: 98.9% of products contain at least one, yet only 24% cover all eight dimensions of the AISPA taxonomy. Third, system prompts have grown steadily longer and more protective of users, suggesting that user protection is becoming a more visible concern in commercial prompt design. Fourth, despite this progress, problematic instructions remain pervasive: roughly 40% of products contain at least one instruction that works against user interests, and protective and problematic instructions frequently coexist within the same prompt. Our findings highlight the need for greater transparency, standardization, and independent oversight for system prompts in commercial AI products.