arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CallBench:电话助手双目标协调基准测试

CallBench: A Benchmark for Dual-Goal Coordination in Phone Call Assistants

Xuzhao Geng, Haozhao Wang, Xuelian Li, Zhenyu Yang, Haonan Lu, Rui Zhang, Ruixuan Li

arXiv 2607.22635首次发表:更新:

AI 中文总结

介绍用于评估电话助手双目标协调的CallBench基准测试,含5万条多轮电话对话及六种场景,设计了评估协议,通过实验表明现有方法在此任务有困难,凸显能在代理约束下做可靠轮次级决策的电话助手的需求。

AI 中文摘要

面向目标的对话系统在通过交互对话完成用户目标方面展现出强大能力。但现有研究主要针对单一明确目标的完成,而电话助手面临代理设置,需协调设备所有者的明确预设目标与来电者的隐含动态目标。我们引入了CallBench,这是一个用于评估电话助手双目标协调的中文基准测试。它包含5万个完整的多轮电话对话,涵盖六种场景,涉及多种预设情况及双方目标的不同关系。还设计了预设感知的轮次级评估协议。实验表明现有方法在此任务上仍有困难,凸显了在代理约束下能在两个独立目标间做出可靠轮次级决策的电话助手的必要性。

英文摘要

Target-oriented dialogue systems have demonstrated strong capabilities in completing user goals through interactive conversations. However, existing studies are primarily designed for single, explicit goal completion, while phone call assistants face a proxy setting that requires coordinating the device owner's explicit preset goal with the caller's implicit and dynamic goal. We introduce \textsc{CallBench}, a Chinese benchmark for evaluating dual-goal coordination in phone call assistants. \textsc{CallBench} contains 50,000 complete multi-turn phone call dialogues across six scenarios: takeout, delivery, taxi, work, life, and harassment. It covers regular presets, emergent presets, and no-preset cases, and includes diverse relations between owner-side and caller-side goals, such as alignment, complementarity, irrelevance, and conflict. We further design a preset-aware turn-level evaluation protocol covering semantic understanding, context use, active guidance, response quality, preset compliance, dialogue rhythm, and safety. Experiments on representative dialogue methods show that existing approaches still struggle with this task, highlighting the need for phone call assistants that can make reliable turn-level decisions between two independent goals under proxy constraints.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑