arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Latent-IM:面向语音大语言模型的潜在交互管理

Latent-IM: Latent Interaction Management for Speech LLMs

Adar Avsian, Atahan Dokme, Tony Woo, Larry Heck

arXiv 2607.26928首次发表:更新:

AI 中文总结

该研究针对语音大语言模型提出Latent-IM框架,将对话动作控制分为选择与实现两个耦合问题,可提升对话动作准确率12.5个百分点,性能与微调相当。

AI 中文摘要

经典口语对话系统常将对话管理与响应生成分离:策略模块选择下一个对话动作,生成组件将该动作表达出来。随着对话系统转向大语言模型(LLM),这种分解已在很大程度上融入模型的隐藏表示中。我们探究是否能恢复LLM内部类似状态估计与动作控制的机制,用于确认、核查、查询、解释、回复等对话动作。我们将动作控制表述为两个耦合问题:选择,即从对话上下文预测合适的下一个动作;实现,即生成时以因果方式生成选定的动作。我们引入Latent-IM,这是一个内部对话管理框架,为在不同目标下选择和部署对话动作提供通用接口。在此,我们利用该控制机制复现人类的动作选择,与未引导的主干模型相比,将平均端到端动作准确率提升了12.5个百分点,同时性能与微调相当。

英文摘要

Classical spoken dialogue systems often separated dialogue management from response realization: a policy selected the next dialogue action, and a generation component expressed that action. As dialogue systems shift toward LLMs, this decomposition has largely disappeared into the model's hidden representations. We ask whether an LLM-internal analogue of state estimation and action control can be recovered for conversational moves such as acknowledging, checking, querying, explaining, and replying. We formulate move control as two coupled problems: selection, predicting the appropriate next move from the dialogue context, and realization, causally producing a chosen move at generation time. We introduce Latent-IM, an internal dialogue-management framework that provides a general interface for choosing and deploying conversational moves under different objectives. Here, we use this control to reproduce human move choices, improving average end-to-end move accuracy by 12.5 points over the unsteered backbone while performing comparably to fine-tuning.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑