发表机构
Nankai University; University of Waterloo; Nanjing University(南开大学; 滑铁卢大学; 南京大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对全模态大语言模型助手式交互的评估瓶颈,构建OmniAssistBench数据集,实验显示Gemini-3-Pro、Qwen3-Omni-Instruct表现存在差距,当前模型在多轮交互等方面仍有不足。
AI 中文摘要
近期,全模态大语言模型(Omni-LLMs)作为实时视频助手展现出巨大潜力,这类模型可持续感知环境并引导用户达成特定目标。与传统被动视频理解不同,交互式助手需主动结合视觉状态、用户目标与先验知识以提供有效帮助。评估这类助手颇具挑战性,因为模型不可预测的响应会动态改变用户后续动作,而静态离线数据集无法适配这种特性。为解决这一瓶颈,我们推出OmniAssistBench。针对同一用户目标可通过多种方法实现导致交互路径发散的问题,我们为模型提供源自源视频的预定义先验,要求其引导用户沿完全相同的路线前进。由于真实交互视频稀缺,我们通过对现有互联网视频进行逆向工程构建该数据集:推导合理的用户目标,将视频分割为多轮片段以模拟连续交互。构建该数据集的严格流程耗费了超过1000个专家工时。实验结果显示,专有模型Gemini-3-Pro得分66.4(满分100),开源模型Qwen3-Omni-Instruct得分51.2。尽管当前模型通常能理解用户输入,但常提供错误或不完整的答案,具体表现为难以处理视觉提示(如手势)、多轮交互中无法维持历史上下文、未能延迟响应至目标事件。结果表明,模型要成为可靠助手仍有巨大改进空间。
英文摘要
Recent omni-modal large language models (Omni-LLMs) show great potential as real-time video assistants, which continuously perceive environments and guide users to achieve specific goals. Unlike traditional passive video understanding, interactive assistants should actively combine visual states, user goals, and prior knowledge to provide effective help. Evaluating this is rather challenging, as the model's unpredictable response dynamically changes the user's subsequent actions, which static offline datasets cannot accommodate. To address this bottleneck, we introduce OmniAssistBench. To solve the issue of diverging interaction paths where the same user goal can be achieved through various methods, we provide models with predefined priors derived from the source video, requiring them to guide users along the exact same routes. Since real interaction videos are rare, we construct the dataset by reverse-engineering existing Internet videos. We deduce logical user goals and segment the videos into multi-turn clips to simulate continuous interactions. This rigorous pipeline required over 1000 expert person-hours to build the dataset. Results show that the proprietary Gemini-3-Pro reaches 66.4 out of the max point of 100, while the open-source Qwen3-Omni-Instruct achieves 51.2. Although current models generally understand user inputs, they frequently provide incorrect or incomplete answers. Specifically, they struggle with visual prompts (e.g., hand gestures), fail to maintain historical context during multi-turn interactions, and fail to delay response until the target event. Results indicate substantial room for improvement before models can become reliable assistants.
CommentsProject page: https://xianyunsun.github.io/OmniAssistBench/