发表机构
Shanghai Qi Zhi Institute; Xiongan AI Institute; Tsinghua University; CUHK-Shenzhen; USTC; Sun Yat-sen University(上海期智研究院; 雄安人工智能研究院; 清华大学; 香港中文大学(深圳); 中国科学技术大学; 中山大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究基于大语言模型的生成式智能体模拟人类行为时中间步骤注释稀缺问题,引入交互式模拟界面收集步骤级人类偏好监督,经监督微调算法和直接偏好优化方法进行步骤级偏好学习,提升模拟效果,证明步骤级人类监督是有效训练信号。
AI 中文摘要
基于大语言模型的生成式智能体通过包括规划、记忆检索、反思和行动选择等中间步骤的长期决策过程来模拟人类行为。然而,对这些中间步骤的细粒度人类注释仍然稀缺,现有智能体并未基于人类对这些中间决策的偏好。为解决这一差距,我们引入了——一种交互式模拟界面,可收集对智能体决策轨迹的步骤级人类偏好监督,从而得到一个包含 57K 细粒度注释的数据集。咱们使用所提方法(\method),通过监督微调算法和直接偏好优化方法,使智能体的社交模拟效果显著提升。实验结果表明,步骤级人类监督是改善局部决策质量和长期智能体行为的有效训练信号。同时这一方法也可以提高模拟的逼真程度、协调能力以及实现更好的交互。
英文摘要
Large language model (LLM)-based generative agents simulate human behavior through long-horizon decision-making processes that comprise intermediate steps such as planning, memory retrieval, reflection, and action selection. However, fine-grained human annotations of these intermediate steps remain scarce, and existing agents are not grounded in human preferences over such intermediate decisions. To address this gap, we introduce \method, an interactive simulation interface that enables us to collect step-level human preference supervision over agent decision trajectories, leading to a dataset of 57K fine-grained annotations. We conduct step-level preference learning on open-weight language models using supervised finetuning and direct preference optimization on this data, consistently improving simulation fidelity, coordination, and interaction quality, and inducing more socially effective agent behavior. Our results show that step-level human supervision is an effective training signal for improving both local decision quality and long-horizon agent behavior.
CommentsWAICA2026