发表机构
PayPal AI(PayPal人工智能部门)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出三层人格向量(含23个维度)用于可控用户模拟,通过64,698个多轮对话验证,证明能产生可测量差异并忠实再现现实难度分布。
AI 中文摘要
评估工具增强的LLM智能体需要多样化、真实的用户输入,然而大多数评估框架使用扁平的角色描述(如“你是一个愤怒的客户”),无论底层场景如何,都会产生几乎相同的对话。在本文中,我们提出了一种具有23个可操作维度的三层人格向量:6个分类人口统计特征(司法管辖区、年龄、渠道、设备、语言熟练度、时间可用性),12个连续行为特质(耐心、自信、数字素养等),这些特质围绕精心设计的档案基向量以高斯噪声采样,以及5个连续情绪状态(沮丧、焦虑、信任、自信、压力),这些状态会随场景上下文而转变。与人格正交,一个4级查询复杂度叠加层控制话语措辞从直接到故意模糊。我们在一个合成数据生成流程中评估该人格模型,涵盖64,698个多轮对话,跨越8个命名档案和3个生产语料库。主要发现:(i)智能体目标达成率在人格间存在15.8个百分点的差距,确认特质向量能产生可测量的不同用户行为;(ii)由于场景反应性情绪状态转变,同一人格在不同场景下表现不同,验证了场景反应性设计;(iii)特定领域项目在预订流程合规性上显示出人格敏感性(分层感知与压力测试人格之间约15-20个百分点的差距),证明该模型忠实再现了现实世界的难度分布;(iv)七条规则描述的特质相关性产生可审计的共现模式,无需学习协方差矩阵。该人格模型已完整指定以供复现。
英文摘要
Evaluating tool-augmented LLM agents requires diverse, realistic user inputs yet most evaluation frameworks use flat role descriptions ("you are an angry customer") that produce near-identical conversations regardless of the underlying scenario. In this paper, we propose a three-tier persona vector with 23 operationalized dimensions: 6 categorical demographics (jurisdiction, age, channel, device, language proficiency, time availability), 12 continuous behavioral traits (patience, assertiveness, digital literacy, etc.) sampled with Gaussian noise around curated profile base vectors, and 5 continuous emotional states (frustration, anxiety, trust, confidence, stress) that shift in response to scenario context. Orthogonal to the persona, a 4-level query-complexity overlay controls utterance phrasing from direct to deliberately vague. We evaluate the persona model inside a synthetic data generation pipeline across 64,698 multi-turn conversations spanning 8 named profiles and 3 production corpora. Key findings: (i) a 15.8 percentage-point spread in agent goal-achievement across personas confirms trait vectors produce measurably different user behavior; (ii) the same persona behaves differently across scenarios due to scenario-reactive emotional state shifts, validating the scenario-reactive design; (iii) domain-specific projects show persona sensitivity on booking-flow compliance (~15-20 percentage points gap between tier-aware and pressure-test personas), demonstrating the model faithfully reproduces real-world difficulty distributions; (iv) seven rule-described trait correlations produce auditable co-occurrence patterns without requiring learned covariance matrices. The persona model is fully specified for reproduction.
Comments7 pages, 3 figures, 6 tables. Extended treatment of the persona component of StateGen (arXiv:2606.16307)