arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

全模态需求理解:多模态交互中上下文用户意图推断的基准

Omni Demand Understanding: A Benchmark for Contextual User-Intent Inference in Multimodal Interaction

Qi Chen, Yunfei Chu, Haolin He, Yifan Yang, Zihan Liu, Yuxuan Wang, Ziyang Ma, Ruiyang Xu, Meng Gao, Yinsong Yan, Ling Wang, Hui Wang, Wen Huang, Yiheng Chen, Guanrou Yang, Qiuqiang Kong, Jin Xu, Xie Chen

arXiv 2609.21392首次发表:更新:

发表机构

Shanghai Jiao Tong University; Shanghai Innovation Institute; Alibaba Token Hub, Alibaba Group; The Chinese University of Hong Kong; Tsinghua University; Hong Kong Polytechnic University; Nankai University; Johns Hopkins University(上海交通大学; 上海创新研究院; 阿里巴巴集团阿里巴巴Token Hub; 香港中文大学; 清华大学; 香港理工大学; 南开大学; 约翰斯·霍普金斯大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对多模态交互中用户需求推断不足的问题,提出ODU基准,通过五维评估和ODU-Bench数据集,发现最强模型仅恢复44.7%关键信息,多数模型误触发率高,揭示系统性能力差距。

AI 中文摘要

自然的音视频交互正成为AI助手的重要界面,使用户能够通过语音和视觉而非精心编写的文本提示进行交流。然而,现有的交互能力基准仍主要关注响应质量,留下了一个更基本的问题未得到充分探索:模型能否从复杂的多模态交互中正确推断用户的潜在需求?现实世界中的用户需求在语音中往往指定不足,必须从多模态线索和对话历史中推断。这种推断因模糊或口吃不清的表达以及嘈杂的声学环境而进一步复杂化。相反,类似请求的语音可能并不构成对助手的需求,导致误触发。我们将全模态需求理解(ODU)确立为一个独特的多模态上下文推断问题:给定一个交互流,模型必须检测用户需求是否存在,并从多模态和对话上下文中推断意图。ODU沿五个维度评估该能力,涵盖单轮和多轮交互。我们使用挑战驱动的分类法、分类法引导的智能体视频生成和人工录制的交互构建了ODU-Bench,随后进行媒体锚定标注和人工验证。我们评估了14个原生MLLM。即使最强的Gemini 3.1 Pro,也只能恢复必须从视觉、声学或对话上下文中推断的关键信息的44.7%。此外,14个模型中有11个在非需求场景中的误触发率超过50%。这些结果揭示了当前MLLM在推断上下文用户需求能力上的系统性差距。我们希望ODU能够建立对多模态交互中一个先前未充分探索但至关重要的能力的评估:在生成适当响应之前正确理解用户需求。

英文摘要

Natural audio-visual interaction is emerging as an important interface for AI assistants, allowing users to communicate through speech and vision rather than carefully composed text prompts. However, existing benchmarks of interactive capabilities still focus primarily on response quality, leaving a more fundamental question underexplored: can a model correctly infer the user's underlying demand from complex multimodal interaction? Real-world user demands are often underspecified in speech and must be inferred from multimodal cues and dialogue history. This inference is further complicated by ambiguous or disfluent expression and noisy acoustic environments. Conversely, request-like speech may not constitute a demand to the assistant, leading to false triggers. We establish Omni Demand Understanding (ODU) as a distinct multimodal contextual inference problem: given an interaction stream, a model must detect whether a user demand is present and infer intent from multimodal and conversational context. ODU evaluates this capability along five dimensions, covering both single-turn and multi-turn interactions. We construct ODU-Bench using a challenge-driven taxonomy, taxonomy-guided agentic video generation, and human-recorded interactions, followed by media-grounded annotation and human verification. We evaluate 14 native MLLMs. Even the strongest, Gemini 3.1 Pro, recovers only 44.7% of key information that must be inferred from visual, acoustic, or conversational context. Moreover, 11 of the 14 models exhibit false-trigger rates above 50% on non-demand scenarios. These results reveal a systematic capability gap in current MLLMs' ability to infer contextual user demands. We hope ODU can establish the evaluation of a previously underexplored yet essential capability in multimodal interaction: correctly understanding user demands before generating an appropriate response.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑