arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.17191cs.AI

迈向拟人化对话:一个用于类人聊天生成、评估和偏好对齐的闭环框架

Toward Anthropomorphic Dialogue: A Closed-Loop Framework for Human-Like Chat Generation, Evaluation, and Preference Alignment

Wentao Liu, Siyu Song, Xi Chen, Youjia Li, Xiaokun Wang, Min Ji, Ji Wang

首次发表
浏览论文内容

中文总结 AI 辅助

该研究提出AnthroDial闭环框架,将拟人化对话作为联合问题,结合角色调度运行时、可执行基准和训练后管道。在多场景基准上评估多个系统,结果显示共享行为维度时拟人化对话有益,不同方法提升了严格准确率。

中文摘要 AI 辅助

类人私人聊天需要的不仅仅是流畅的回复生成:一个系统必须保留角色、关系、记忆、有限知识、特定媒介的时机以及连贯的多轮对话弧线。我们提出了AnthroDial,一个闭环框架,将拟人化对话表述为系统架构、可执行评估和诊断对齐的联合问题。它结合了:(1)一个基于角色的调度对话运行时,带有角色和场景卡片、长期记忆、虚拟时间和单稿消息决策;(2)一个可执行基准,有L0有效性门、五个每轮维度和五个对话级维度;(3)一个训练后管道,为监督微调过滤出16436个调度决策示例,并应用带有认知诊断、最近发展区感知奖励的广义信赖域政策优化。该奖励维护每个行为维度的卡尔曼滤波能力估计,对能力 deficit 较大的维度进行加权,并使用展开分数作为任务级最近发展区匹配,以将优化重点放在可学习的弱技能上。在一个有55个角色、50个场景、50个角色 - 场景绑定以及每个模型100个基于角色的案例的基准上,我们评估了16个系统,涵盖前沿基线、开放模型、思维/无思维变体以及监督微调/强化学习消融。最强的未训练基线达到32.00%的严格准确率,而Qwen3.6 - 27B - SFT + RL达到39.00%的严格准确率和98.5的总分。在9B无思维设置中,监督微调将严格准确率从0.00%提高到13.00%,强化学习提高到18.37%。这些结果表明,当生成、评估和奖励塑造共享相同的行为维度时,可以实现拟人化对话。

英文摘要

Human-like private chat requires more than fluent response generation: a system must preserve persona, relationship, memory, bounded knowledge, medium-specific timing, and a coherent multi-turn arc. We present AnthroDial, a closed-loop framework that formulates anthropomorphic dialogue as a joint problem of system architecture, executable evaluation, and diagnostic alignment. It combines (1) a role-conditioned scheduled dialogue runtime with persona and scenario cards, long-term memory, virtual time, and single-draft message decisions; (2) an executable benchmark with an L0 validity gate, five per-turn dimensions, and five dialogue-level dimensions; and (3) a post-training pipeline that filters 16,436 scheduled-decision examples for SFT and applies GRPO with a cognitive-diagnostic, ZPD-aware reward. The reward maintains Kalman-filtered capability estimates for each behavioral dimension, upweights dimensions with larger capability deficits, and uses rollout scores as task-level ZPD matches to focus optimization on learnable weak skills. On a benchmark with 55 personas, 50 scenarios, 50 persona-scenario bindings, and 100 role-conditioned cases per model, we evaluate 16 systems spanning frontier baselines, open models, thinking/no-think variants, and SFT/RL ablations. The strongest non-trained baseline reaches 32.00% strict ACC, while Qwen3.6-27B-SFT+RL reaches 39.00% strict ACC and a 98.5 overall score. In the 9B no-think setting, SFT and RL improve strict ACC from 0.00% to 13.00% and 18.37%. These results show that anthropomorphic dialogue benefits when generation, evaluation, and reward shaping share the same behavioral dimensions.

发表机构

  • Shanghai Institute of Innovation(上海创新研究院)
  • East China Normal University(华东师范大学)
  • University of Science and Technology of China(中国科学技术大学)
  • Chabiyue (Shanghai) Information Technology Co., Ltd.(上海叉叉悦信息技术有限公司)
  • Shanghai Tianyou Software Co., Ltd.(上海天游软件有限公司)
  • Zhejiang Century Huatong Group Co., Ltd.(浙江世纪华通集团有限公司)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑