评估大语言模型个性化的隐性成本
Evaluating the Hidden Costs of Personalization in Large Language Models
浏览论文内容
中文总结 AI 辅助
本研究针对大语言模型个性化的三种隐性风险,提出PRISK评估框架,经13个模型实证分析,发现个性化信息会加剧偏见并导致特定指标下降。
中文摘要 AI 辅助
尽管大语言模型(LLMs)会整合用户个性化信号以提升可用性和实用性,但在对话历史、推断偏好、用户画像等个性化语境下,它们越来越偏离提供平衡、信息丰富的回复,转而优化用户满意度。具体而言,我们识别出三种新兴风险:(1)无关个性化:模型在不必要的语境中提及个人信息;(2)偏好收窄:模型强化信息回音室;(3)谄媚偏见:模型过度赞同用户观点。因此,模型可能在非必要语境中提及个人信息,无意间降低回复多样性,或过度赞同用户观点。尽管AI助手对个性化的应用日益广泛,但对其潜在副作用的系统评估却十分有限。为填补这一空白,我们提出PRISK,一个包含自动化数据生成和定制化指标的动态评估框架,可揭示当前LLM个性化的系统性局限,以及个性化信息如何影响其回复。我们对13个LLMs的实证分析表明,用户画像和检索记忆会持续加剧偏见,导致无关个性化平均下降45.9%,偏好收窄平均下降41.7%,谄媚偏见平均下降61.7%。
英文摘要
While Large language models (LLMs) incorporate user personalization signals to improve usability and helpfulness, they increasingly shift from providing balanced, informative responses toward optimizing for user satisfaction when conditioned on personal context such as conversation history, inferred preferences, and user profiles. Specifically, we identify three emerging risks: (1) irrelevant personalization, where models reference personal information in unnecessary contexts; (2) preference narrowing, where models reinforce informational echo chambers; and (3) sycophantic bias, where models agree excessively with user opinions. As a result, models may reference personal information in contexts where it is unnecessary, inadvertently collapse response diversity, or agree excessively with user opinions. Despite the growing use of personalization in AI assistants, there has been limited systematic evaluation of its potential side effects. To bridge this gap, we propose PRISK, a dynamic evaluation framework with automated data generation and tailored metrics that uncovers systematic limitations in current LLM personalization and how personalized information shapes its responses. Our empirical analysis across 13 LLMs demonstrates the presence of user profiles and retrieved memories consistently exacerbates biases, resulting in an average drop of 45.9% in irrelevant personalization, 41.7% in preference narrowing and 61.7% in sycophantic bias.
发表机构
- University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
- HKUST(香港科技大学)
机构由 AI 辅助整理,请以论文原文为准。