发表机构
SenseTime Research; University of Science and Technology of China; Tsinghua University(商汤科技研究院; 中国科学技术大学; 清华大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出EYT-Bench基准,通过三方解耦设计评估大语言模型多轮对话能力,引入新指标,发现先进模型在主观维度相近但客观意图跟踪差异大,推理可提升客观跟踪,角色格式影响轨迹分布,且热身效应稳健。
AI 中文摘要
评估大语言模型(LLMs)作为多轮对话伙伴需要单轮基准测试所没有的探测能力:角色一致性、不断演变的意图跟踪、情感动态和目标完成情况。我们引入了EYT-Bench,这是一个围绕三方解耦设计构建的以人为本的基准:一个基于角色的用户模拟器、一个将意图感知与响应生成分开的目标模型,以及一个独立的第三方LLM评判器,可选择多评判器集成。角色是从公共人工策划语料库Nemotron-Personas-USA和PersonaMem-v2中采样,而不是合成的,减少了LLM引起的角色偏差。EYT-Bench还引入了两个轨迹级指标:基于嵌入的意图漂移和最终意图完成率(FICR),灵感来自tau-bench。在17个目标x200个对话的评估中,EYT-Bench揭示了四个发现:(i)在主观维度上,最先进的闭源和开源模型在统计上接近(同理心/角色/拟人化差异<=0.3),但在客观意图跟踪上相差高达9倍;(ii)推理(“思考”)显著提高了长上下文角色的客观跟踪(在Gemma-4上潜在意图准确率提高0.47-0.50),而主观分数几乎不变;(iii)角色格式主导轨迹分布,在Nemotron-USA上FICR在0.95以上饱和,但在PersonaMem-v2上从0.53扩展到0.88;(iv)热身效应在16/17个模型上很稳健(一个异常值GPT-5.5逆转了效应),在[0.05,0.15]范围内的alpha上排名稳定。使用deepseek-v4-pro进行的交叉评判消融证实,目标排名和最终意图满意度在评判器之间是保持不变的。
英文摘要
Evaluating large language models (LLMs) as multi-turn conversational partners requires probing capabilities that single-turn benchmarks miss: persona consistency, evolving intent tracking, emotional dynamics, and goal completion across many turns. We introduce EYT-Bench, a human-centered benchmark whose evaluation protocol is built around a decoupled three-party design: a persona-grounded user simulator, a target model evaluated on both intent perception and response generation, and an independent, configurable ensemble of LLM judges. Across 3,400 dialogues with 17 target models, EYT-Bench reveals four findings that previous benchmarks miss: (i) state-of-the-art closed and open-source models are statistically indistinguishable on subjective dimensions, but separate by up to 9x on objective intent-tracking; (ii) reasoning is a phase transition for objective tracking on long-context personas but is essentially flat on subjective scores; (iii) persona format strongly affects trajectory spread, FICR (final-intent completion rate) saturates above 0.95 on Nemotron-USA but ranges from 0.53 to 0.88 on PersonaMem-v2; and (iv) the warm-up effect is observed in 16 of 17 models.