arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.28378cs.CL

PersonaForge:面向智能体系统的逼真多轮用户模拟

PersonaForge: Realistic Multi-Turn User Simulation for Agentic Systems

发表机构北京大学 · 小米 · 香港大学
另 1 家 · 查看机构详情
  • Peking University(北京大学)
  • Xiaomi(小米)
  • The University of Hong Kong(香港大学)
  • Renmin University of China(中国人民大学)

机构由 AI 辅助整理,请以论文原文为准。

Hanglong Lv, Dawei Zhu, Lei Li, Bowen Ye, Huaqiu Liu, Yifan Song, Bofei Gao, Weimin Xiong, Jinhao Dong, Chenhong He, Lingpeng Kong, Qi Liu, Tong Yang, Fuli Luo

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对现有智能体系统训练评估中多轮用户交互数据不足的问题,提出PersonaForge模拟框架构建相关数据集与基准,实验证实其可提升智能体交互效率与任务表现。

中文摘要 AI 辅助

大型语言模型正越来越多地被用作智能体工作流执行器,但现有的训练数据和基准大多假设查询信息完整且为单轮。我们对1.6万条真实会话的分析显示,75.9%的交互属于多轮,这表明用户与智能体的交互方式与这类系统的训练、评估方式之间存在巨大差距。我们推出PersonaForge,这是一个用于合成逼真多轮用户-智能体交互的用户模拟框架。PersonaForge结合了四维角色空间、基于真实用户统计数据校准的SOUL驱动行为控制,以及基于真实种子查询的反向深度构建。利用PersonaForge,我们构建了一个包含6300条记录的训练数据集,以及PersonaForge-Bench——一个手动标注的138项任务基准,涵盖20多个专业领域并采用四维评分。对Qwen3.5-27B的实验表明,PersonaForge训练使综合评分提升了4.1%,且在四个维度均有提升,其中任务完成度(+6.0%)和响应质量(+6.8%)的提升最大。进一步分析显示,经PersonaForge训练的智能体使用的轮次和工具调用更少,表明交互效率得到提升,而消融实验证实了SOUL组件和自适应模拟的贡献。总体而言,PersonaForge和PersonaForge-Bench为在逼真多轮用户交互场景下训练和评估智能体奠定了基础。

英文摘要

Large language models are increasingly used as agentic workflow executors, yet existing training data and benchmarks largely assume informationally complete, single-turn queries. Our analysis of 16K real-world sessions shows that 75.9% of interactions are multi-turn, revealing a substantial gap between how users interact with agents and how such systems are trained and evaluated. We introduce \textbf{PersonaForge}, a user simulation framework for synthesizing realistic multi-turn user--agent interactions. PersonaForge combines a four-dimensional persona space, SOUL-driven behavioral control calibrated to real-user statistics, and Reverse Deep Construction grounded in authentic seed queries. Using PersonaForge, we construct a 6.3K-record training dataset and \textbf{PersonaForge-Bench}, a manually annotated 138-task benchmark spanning over 20 professional domains with four-dimensional scoring. Experiments on Qwen3.5-27B show that PersonaForge training improves the composite score by +4.1%, with gains across all four dimensions and the largest improvements in Task Completion (+6.0%) and Response Quality (+6.8%). Further analyses show that PersonaForge-trained agents use fewer turns and tool calls, suggesting improved interaction efficiency, while ablations confirm the contribution of SOUL components and adaptive simulation. Together, PersonaForge and PersonaForge-Bench establish a foundation for training and evaluating agents under realistic multi-turn user interaction.

相关深度报道

↑