arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.37853cs.CL

AnthroDial:自主社交互动中LLM拟人化的基准测试

AnthroDial: Benchmarking LLM Anthropomorphism in Autonomous Social Interaction

Wentao Liu, Xi Chen, Siyu Song, Biao Yuan, Yu Zhang, Zhou Zhuotong, Jingying Zhou, Guohao Feng, Shasha Hu, Tianfu Wang, Shangshang Yang, Haoyang Liu, Youjia Li,… 展开作者

Wentao Liu, Xi Chen, Siyu Song, Biao Yuan, Yu Zhang, Zhou Zhuotong, Jingying Zhou, Guohao Feng, Shasha Hu, Tianfu Wang, Shangshang Yang, Haoyang Liu, Youjia Li, Xiaokun Wang, Min Ji, Ji Wang

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出AnthroDial框架,通过MindFlow、CAPS-Eval和训练范式三方面提升LLM社交智能体的拟人化自主互动能力,实验验证了其有效性和评估可靠性。

中文摘要 AI 辅助

大型语言模型(LLMs)日益被部署为社交智能体,然而可信的类人互动需要的不仅仅是流畅的回应或角色一致性。智能体必须自主决定是否、何时以及如何沟通,同时适应不断变化的背景、目标和关系。然而,现有研究缺乏一种统一的方法来在持续、开放式的互动中启用、评估和改进此类能力。我们引入了AnthroDial,一个从三个互补方面开发拟人化社交智能体的统一框架:MindFlow,一个轻量级交互框架,通过动态思维缓冲区实现自主、异步和自适应的沟通;CAPS-Eval,一个基于理论的评估框架,用于评估拟人化互动的认知、情感和行为维度;以及一个可扩展的训练范式,结合了用于环境扩展的SEEDS和用于自适应能力优化的DiAPO。我们进一步构建了涵盖日常交流、游戏互动和长时程角色互动的评估数据集。跨多种模型和场景的大量实验证明了互动自主性和自然性的提升,验证了CAPS-Eval的可靠性、区分性以及与人类排名的一致性,并确认了我们训练范式的有效性。总之,这些组件为在开放式互动中开发可信的类人社交智能体提供了一个统一框架。

英文摘要

Large language models (LLMs) are increasingly deployed as social agents, yet credible human-like interaction requires more than fluent responses or persona consistency. Agents must autonomously decide whether, when, and how to communicate while adapting to evolving contexts, goals, and relationships. Existing research, however, lacks a unified approach to enabling, evaluating, and improving such capabilities in continuous, open-ended interaction. We introduce AnthroDial, a unified framework for developing anthropomorphic social agents from three complementary aspects: MindFlow, a lightweight interaction harness that enables autonomous, asynchronous, and adaptive communication through a dynamic Mind Buffer; CAPS-Eval, a theory-grounded framework for evaluating cognitive, affective, and behavioral dimensions of anthropomorphic interaction; and a scalable training paradigm that combines SEEDS for environment expansion with DiAPO for adaptive capability optimization. We further construct evaluation datasets covering everyday communication, game interaction, and long-horizon character interaction. Extensive experiments across diverse models and scenarios demonstrate improved interaction autonomy and naturalness, validate the reliability, discriminativeness, and agreement with human rankings of CAPS-Eval, and confirm the effectiveness of our training paradigm. Together, these components provide a unified framework for developing credible human-like social agents in open-ended interaction.

发表机构

  • Shanghai Institute of Innovation(上海创新研究院)
  • University of Science and Technology of China(中国科学技术大学)
  • East China Normal University(华东师范大学)
  • Anhui University(安徽大学)
  • Shanghai Jiaotong University(上海交通大学)
  • Fudan University(复旦大学)
  • University of Melbourne(墨尔本大学)
  • Zhejiang University(浙江大学)
  • The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
  • Shanghai Tianyou Software Co., Ltd.(上海天游软件有限公司)
  • Chabiyue (Shanghai) Information Technology Co., Ltd.(茶百悦(上海)信息技术有限公司)
  • Zhejiang Century Huatong Group Co., Ltd.(浙江世纪华通集团股份有限公司)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑