arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.29196cs.CL

Hy-MultiTurn:用于深度多轮对话理解的六维度基准

Hy-MultiTurn: A Six-Dimensional Benchmark for Deep Multi-Turn Dialogue Understanding

Eileen Ye, Jiawen Tao, Yaoming Li, Chenxu Liu, Wenhan Yu, Yaxin Fan, Xiaokun Yuan, Mengzhou Wu, Yanbing Jiang, Maxm Pan

首次发表
浏览论文内容

中文总结 AI 辅助

Hy-MultiTurn是含六种评估模式的中文深度多轮对话理解基准,含209个长对话任务,评估22种前沿模型发现其具挑战性,GPT-5.5仅41.1%响应达标且无模型全模式最优

中文摘要 AI 辅助

与聊天机器人和智能体的长期多轮交互现已十分普遍,正确响应通常依赖于记住早期细节、跟踪后续修订、识别意图对象或指称项,以及在未满足所需条件时弃权(不执行)。现有多轮基准通常覆盖简短交互,无法充分评估长多轮交互中的这些能力,尤其是中文场景下的相关能力,且对模型失败的方式和原因提供的见解有限。为解决这些局限,我们分析真实聊天机器人失败案例以识别六种反复出现的机制,并将其用于定义 Hy-MultiTurn(一个用于深度多轮对话理解的中文基准)中的六种受控评估模式。这六种模式分别评估约束记忆、精确执行、约束综合、对象定位、动作抑制和指称消解。我们在六种模式下构建了209个受控任务,对话轮数跨度为12至76轮,对话长度、无关主题干扰和口语化表述进一步增加了难度。对22种前沿模型配置的评估表明,Hy-MultiTurn具有广泛的挑战性,因为即使是整体表现最强的配置GPT-5.5,也仅有41.1%的响应满足所有要求,且没有任何模型在全部六种模式中表现最佳。

英文摘要

Long-running multi-turn interactions with chatbots and agents are now common, and a correct response often depends on remembering earlier details, tracking later revisions, identifying intended objects or referents, and withholding action when required conditions are unmet. Existing multi-turn benchmarks typically cover short exchanges and do not fully evaluate these capabilities in long multi-turn interactions, particularly in Chinese, while offering limited insight into how and why models fail. To address these limitations, we analyze real chatbot failures to identify six recurring mechanisms and use them to define six controlled evaluation modes in Hy-MultiTurn, a Chinese benchmark for deep multi-turn dialogue understanding. The six modes evaluate constraint memory, precise execution, constraint synthesis, object localization, action suppression, and reference resolution. Across the six modes, we construct 209 controlled tasks spanning 12-76 turns, with dialogue length, irrelevant-topic distraction, and colloquial phrasing adding further difficulty. Evaluation of 22 frontier model configurations shows that Hy-MultiTurn is broadly challenging, as even GPT-5.5, the strongest overall configuration, satisfies all requirements in only 41.1 percent of responses and no model performs best in all six modes.

发表机构

  • Tencent(腾讯)
  • Peking University(北京大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑