识别个性化生成式AI系统中的危害需要在交互层面开展以用户为中心的审计
Identifying Harm in Personalized, Generative AI Systems Requires User-Centered Auditing at the Interaction Level
- Microsoft Research(微软研究院)
- Stanford University(斯坦福大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文指出现有审计方法无法捕捉个性化生成式AI系统的新兴危害,提出应将危害理解重构为自适应的以用户和社区为中心的过程,将审计转向支持交互中持续明确危害的基础设施。
AI中文摘要:
个性化生成式AI系统会随时间推移不断调整自身行为以适配单个用户,从根本上改变了模型的表现。尽管现有的审计方法在揭露非个性化场景中的危害方面已卓有成效,但它们往往依赖静态的模拟评估,以及在宽泛的群体类别中汇总得出的危害定义。在这篇立场文件中,我们认为此类方法无法捕捉个性化生成式AI系统中出现的新兴危害,这类危害通过对持续交互的解读得以显现,并随用户历史而演变。我们确定了诸多危害审计范式背后的三个预设:即危害可(1)在现实交互之外被明确规定,(2)在群体内部以非多元方式被定义,(3)被视为静态。有人可能会辩称,个性化系统只需通过反复交互就能学习到危害对单个用户而言的定义。然而,我们认为,试图通过更深层次的个性化来揭露用户危害的做法,可能会给边缘化用户带来不对称的劳动负担和隐私负担。因此,我们提议将对危害的理解重新构建为自适应的、以用户和社区为中心的过程,并概述了将审计从回顾性评估转向支持在交互中持续明确危害的基础设施的设计方向。我们的工作强调,需要有审计和设计实践,以更好地反映个性化生成式AI系统中危害理解的多元性和演变性。
英文摘要:
Personalized, generative AI systems increasingly adapt their behavior to individual users over time, fundamentally changing model behavior. While existing auditing approaches have been effective at surfacing harms in non-personalized contexts, they often rely on static, simulated evaluations and definitions of harm that aggregate across broad, group categories. In this position paper, we argue that such approaches can fail to capture emergent harms in personalized generative AI systems, where harms surface through interpretations of ongoing interaction and evolve with user history. We identify three presuppositions underlying many harm auditing paradigms: that harms can be (1) specified outside real-world interaction, (2) defined non-pluralistically within groups, and (3) treated as static. One might argue that personalized systems could simply learn definitions of what constitutes harm to individual users through repeated interactions. However, we argue that attempts to surface user harms through deeper personalization risk imposing asymmetric burdens of labor and privacy on marginalized users. Consequently, we propose reframing understandings of harm as adaptive, user- and community-centered processes, and outline design directions that shift auditing from retrospective evaluation toward infrastructures that support ongoing articulation of harm in interaction. Our work highlights the need for auditing and design practices that better reflect the pluralistic and evolving nature of harm understanding in personalized generative AI systems.