arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.00921cs.AIcs.CL

VIBE-Bench:当用户画像不代表偏好时评估个性化大语言模型

VIBE-Bench: Evaluating Personalized Large Language Models When Profiles Don't Mean Preferences

  • Monash University(莫纳什大学)
  • Singapore Management University(新加坡管理大学)
  • University of Liverpool(利物浦大学)
  • RMIT University(皇家墨尔本理工大学)

机构由 AI 辅助整理,请以论文原文为准。

Yiwen Jiang, Yang Deng, Stephanie Fong, Zimu Wang, Yaling Shen, Wei Feng, Hongxi Yang, Xiangyu Zhao, Zhongxing Xu, Deval Mehta, Xuelian Cheng, Zongyuan Ge

AI总结:

本文提出VIBE-Bench基准,针对用户画像与偏好概念错位的场景,发现现有个性化大语言模型依赖浅层语义关联,无法完成跨概念偏好推理,该基准可用于推进相关研究。

AI中文摘要:

个性化大语言模型(PLLMs)旨在为不同用户定制响应,核心挑战是偏好推理:从用户相关历史中推断与查询相关的偏好。然而,现有基准大多假设可从语义相关的历史中检索此类偏好。本文研究一种未被充分探索但具有实际重要性的场景——画像-偏好概念错位(PRCM),其中可观测的画像线索与查询特定偏好处于不同概念空间,导致语义检索无法适配个性化需求。我们推出VIBE-Bench,这一基准包含两项基于心理学的任务、3504个用户画像、12239段对话,还包含经人工验证的黄金测试集,要求模型进行超越表面语义重叠的跨概念偏好推理。对多种个性化方法的实验表明,当前PLLMs大多依赖浅层语义关联,无法习得稳健的跨概念映射。这些发现将PRCM确立为PLLMs中一种独特的失效场景,并将VIBE-Bench定位为推进超越语义匹配的偏好推理的专用测试平台。

英文摘要:

Personalized Large Language Models (PLLMs) aim to tailor responses to individual users, where a central challenge is preference reasoning: inferring query-relevant preferences from user-related history. Existing benchmarks, however, largely assume that such preference can be retrieved from semantically related history. We study an underexplored but practically important regime, profile-preference conceptual misalignment (PRCM), where observable profile cues and query-specific preferences lie in different concept spaces, making semantic retrieval inconsistent for personalization. We introduce VIBE-Bench, a benchmark with two psychology-grounded tasks, 3,504 personas and 12,239 dialogues, including a manually verified gold test set, and requires cross-concept preference reasoning beyond surface semantic overlap. Experiments with several personalization methods show that current PLLMs largely rely on shallow semantic correlations and fail to acquire robust cross-concept mappings. These findings establish PRCM as a distinct failure regime in PLLMs and position VIBE-Bench as a focused testbed for advancing preference reasoning beyond semantic matching.

补充信息

↑