AgentPersonaBench:基准测试人格驱动的用户模拟
AgentPersonaBench: Benchmarking Persona-Driven User Simulation
浏览论文内容
中文总结 AI 辅助
AgentPersonaBench通过四个交互界面评估人格驱动用户模拟的行为保真度,发现领先模型可达到84.7%的遵循率,但跨模态和多属性场景存在明显局限。
中文摘要 AI 辅助
我们提出了AgentPersonaBench(APB),这是一个用于评估人格条件作用是否忠实地引导下游智能体行为的基准。尽管语言模型越来越多地被部署用于人格驱动的用户模拟,但现有基准主要评估对话风格或自我报告,而非真实的行为保真度。APB每次评估一个潜在人格特质,将每个目标特质嵌入到完整的合成档案中,而不显式命名该特质或披露测试。真实性的遵循度严格通过可观察行为在四个日益逼真的交互界面(调查、聊天、网页(交互式网页环境)和应用(桌面软件环境))中进行验证。APB包含2,460个任务,涵盖867个特质,并通过自动化审计和专家评审进行验证。我们对20个前沿模型臂的评估表明,高保真用户模拟已经可以实现:领先模型在无提示条件下达到高达84.7%的完全通过遵循率。同时,APB识别出清晰的行为边界:遵循率随交互模态下降(仅37.9%-64.3%通过所有四个界面),多属性需求降低保持率,且竞争模型家族表现出显著的行为差异。
英文摘要
We introduce AgentPersonaBench (APB), a benchmark evaluating whether persona conditioning faithfully steers downstream agent behavior. While language models are increasingly deployed for persona-driven user simulation, existing benchmarks primarily evaluate conversational styling or self-reports rather than authentic behavioral fidelity. APB evaluates latent persona adherence one trait at a time, embedding each target trait within a complete synthetic profile without explicitly naming the trait or disclosing the test. Ground-truth adherence is verified strictly from observable actions across four interaction surfaces of increasing realism: survey, chat, web (interactive web environments), and app (desktop software environments). APB comprises 2,460 tasks spanning 867 traits, verified through automated audits and expert review. Our evaluation of 20 frontier model arms demonstrates that high-fidelity user simulation is already attainable: leading models achieve up to 84.7% full-pass adherence under unprompted conditions. At the same time, APB identifies clear behavioral boundaries: adherence drops across interaction modalities (only 37.9-64.3% pass all four surfaces), multi-attribute demands degrade retention, and competing model families exhibit pronounced behavioral divergence.