发表机构
KAIST; University of Michigan; Seoul National University; ETH Zürich; SkillBench(韩国科学技术院; 密歇根大学; 首尔大学; 苏黎世联邦理工学院; SkillBench)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对LLM用户模拟器的助手偏差,通过对比对话的用户与助手视角提取用户角色向量,证实偏差结构可识别且用户行为可定向分析,为优化模拟器提供了表征层面依据。
AI 中文摘要
基于大语言模型(LLM)的用户模拟器正被越来越多地用于大规模评估自主智能体,以替代成本高昂的人工评估。尽管这类模拟器前景广阔,但它们存在“助手偏差”——倾向于合作并追求任务目标,极少复现真实用户表现出的沮丧或脱离状态,损害了评估的有效性。此前研究指出,这种偏差在模型训练时就已形成,角色扮演提示无法覆盖它。我们从模型激活角度分析该偏差,通过对比模型对同一对话的用户视角与助手视角的表征,提取用户角色向量。我们得到两项发现:(i)用户方向可在激活中识别,能引发类用户行为,且具有与助手特质不同的特征;(ii)尽管用户角色激活与模拟真实性相关,调整可增强这种相关性,但它可能过度放大用户行为并覆盖个体用户特征。综上,我们的发现为LLM用户模拟器提供了表征层面的分析,确认助手偏差在结构上可识别,且用户行为可进行定向分析。
英文摘要
LLM-based user simulators are increasingly used to evaluate autonomous agents at scale, in place of costly human evaluations. Despite this promise, these simulators exhibit "assistant bias," a tendency to cooperate and pursue task goals. They rarely reproduce the frustration or disengagement that real users exhibit, compromising evaluation validity. Prior work outlines that this bias is baked in during model training, which role-playing prompts fail to override. We analyze this bias from model activations, extracting a user role vector by contrasting how the model represents user versus assistant perspectives on the same dialogue. We observe two findings: (i) the user direction is identifiable in activations, elicits user-like behaviors, and captures characteristics distinct from assistant traits; and (ii) although user-role activation associates with simulation realism and steering strengthens it, it can exaggerate user behaviors and override individual user profiles. Together, our findings provide a representation-level analysis of LLM user simulators, confirming that assistant bias is structurally identifiable and that user behavior can be directionally analyzed.
Comments35 pages, EMNLP 2026 Findings