arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Hear2Act:对韵律应何时改变助手行为的基准测试

Hear2Act: Benchmarking When Prosody Should Change What an Assistant Does

Xinyi Liu, Hooshang Nayyeri, Dilek Hakkani-Tur, Emine Yilmaz, Joo-Kyung Kim, Yifei Zhang, Charith Peris, Hari Thadakamalla

arXiv 2608.19515首次发表:更新:

发表机构

Amazon; University of Illinois Urbana-Champaign; University College London(亚马逊公司; 伊利诺伊大学厄巴纳-香槟分校; 伦敦大学学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究推出 Hear2Act 基准,评估发现词汇证据不足时韵律对助手决策很重要,具备音频能力的 LLM 能从语音恢复信息但需显式中间表示才能可靠转化为行动。

AI 中文摘要

韵律线索能够传达与任务相关的信息,即便用词本身未变,也会改变面向任务的对话的轨迹和结果。然而现有基准通常孤立地评估韵律感知、响应恰当性和面向任务的对话,难以测试韵律证据是否会改变下游决策。我们推出 Hear2Act,这是一个针对文本和语音助手的统一评估协议,包含 480 个基于角色的场景、隐藏的用户关注点以及可客观验证的结果。对于每个场景,我们在固定任务和用户需求的同时,改变同一关注点是通过词语明确传达还是主要通过韵律传达,并在文本、音频和关注点状态访问三种条件下评估决策。利用 Hear2Act,我们评估了两个具备音频能力的大语言模型(LLM)。在韵律介导反馈下,向文本添加音频仅使平均最优解决方案率从 14.6% 变为 15.3%。相比之下,当模型从音频中推断关注点状态、将其表示为文本并用于下一个动作选择时,该比率上升至 39.6%,接近真实状态下的 40.7%。不过,当关注点在话语中被口头提及(即显式词汇反馈)时,这种对比在很大程度上消失了。综合来看,这些结果表明,当词汇证据不足时,韵律很重要;具备音频能力的 LLM 能够从语音中恢复信息,但如果没有显式的中间表示,就无法可靠地将其转化为行动。

英文摘要

Prosodic cues can convey task-relevant information that alters the trajectory and outcome of a task-oriented dialogue, even when the words themselves remain unchanged. Yet existing benchmarks typically evaluate prosodic perception, response appropriateness, and task-oriented dialogue in isolation, making it difficult to test whether prosodic evidence changes downstream decisions. We introduce Hear2Act, a unified evaluation protocol for text and spoken assistants with 480 persona-grounded scenarios, hidden user concerns, and objectively verifiable outcomes. For each scenario, we keep the task and user needs fixed while varying whether the same concern is conveyed explicitly in words or primarily through prosody, and evaluate decisions under transcript, audio, and concern-state access. Using Hear2Act, we evaluate two audio-capable LLMs. Under Prosody-mediated feedback, adding audio to the transcript changes the average optimal-solution rate only from 14.6% to 15.3%. In contrast, when models infer the concern status from audio, represent it in text, and use it for next-action selection, the rate rises to 39.6%, close to 40.7% with the ground-truth state. This contrast, however, largely disappears under Explicit lexical feedback, where the concern is verbally mentioned in the utterance. Together, these results show that prosody matters when lexical evidence is insufficient, and that audio-capable LLMs can recover information from speech but do not reliably carry it into action without an explicit intermediate representation.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑