发表机构
Shenzhen University of Advanced Technology; Ant Group(深圳理工大学; 蚂蚁集团)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究推出首个基于衣食住行等通用领域用户行为日志的LUNAR基准,经实验揭示LLM个性化的关键影响因素,为个性化LLM的发展指明了方向。
AI 中文摘要
现有的个性化大语言模型(LLM)基准主要依赖文本人设或孤立的行为信号,对跨领域行为个性化的评估有限,而跨领域行为个性化要求响应必须基于异质的日常生活活动。为解决这一缺口,我们推出LUNAR,这是首个用于评估LLM如何基于涵盖衣食住行等通用日常生活领域的纵向应用交互历史来个性化响应的基准。为支持可扩展的基准构建,同时缓解数据稀疏性和隐私问题,LUNAR采用基于真实行为模式的多阶段粗到细合成流程。保真度分析显示,LUNAR与真实行为分布的契合度优于其他合成基准。对19种主流LLM的实验表明,获取行为日志是深度个性化的必要条件,但非充分条件:更多上下文或更大模型无法保证更好性能;有效个性化取决于跨领域选择和整合相关证据。直接检索细粒度行为记录始终优于压缩记忆,而更强的个性化可能以隐私保护为代价。这些发现确定了证据选择、跨领域整合和隐私控制是个性化LLM的关键挑战。
英文摘要
Existing personalized LLM benchmarks primarily rely on textual personas or isolated behavioral signals, providing limited evaluation of cross-domain behavioral personalization, where responses must be grounded in heterogeneous daily-life activities. To address this gap, we introduce LUNAR, the first benchmark for evaluating how LLMs personalize responses from longitudinal app interaction histories across universal daily-life domains, including clothing, food, housing, and mobility. To support scalable benchmark construction while mitigating data sparsity and privacy concerns, LUNAR uses a multi-stage coarse-to-fine synthesis pipeline grounded in real-world behavioral patterns. Fidelity analyses show closer alignment with real behavioral distributions than other synthetic benchmarks. Experiments on 19 mainstream LLMs show that access to behavioral logs is necessary but not sufficient for deep personalization: neither more context nor larger models guarantees better performance; effective personalization depends on selecting and integrating relevant evidence across domains. Direct retrieval of fine-grained behavioral records consistently outperforms compressed memory, while stronger personalization can come at the cost of privacy protection. These findings identify evidence selection, cross-domain integration, and privacy control as key challenges for personalized LLMs.