发表机构
School of Information Science and Electronic Engineering, Shanghai Jiao Tong University; Ant Group(上海交通大学信息科学与工程学院; 蚂蚁集团)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对实时视频理解中低延迟交互的缺失,提出统一边缘-云系统,集成六种VLM后端,经WebRTC路径实现约0.9-1.0秒首文本和1.3-1.5秒首音频的低延迟响应。
AI 中文摘要
视觉语言模型正在将视频理解从离线片段分析扩展到连续交互式流媒体,但大多数研究仍强调模型能力而非可部署的低延迟交互。本文提出一个统一的边缘-云系统,用于实时视频VLM应用。轻量级手机、智能眼镜、PC和伪重放客户端将视频和语音发布到服务器运行时,该运行时提供共享的ASR/TTS、会话编排、后端适配、响应交付和基于存档的测量。该系统集成了六个具有流式或交互导向能力的代表性视频VLM后端,并在后端运行时、媒体传输、客户端观察到的延迟和交互行为方面对它们进行评估。通过合适的后端选择和WebRTC路径,测试系统达到约0.9至1.0秒的首个VLM文本输出,以及1.3至1.5秒的首个非静音TTS音频输出,同时暴露了后端适配成本和实时交互行为的差异。
英文摘要
Vision-language models are extending video understanding from offline clip analysis to continuous interactive streaming, but most research still emphasizes model capability rather than deployable low-latency interaction. This paper presents a unified edge-cloud system for real-time video VLM applications. Lightweight phone, smart glasses, PC, and pseudo-replay clients publish video and speech to a server runtime that provides shared ASR/TTS, session orchestration, backend adaptation, response delivery, and archive-backed measurement. The system integrates six representative video VLM backends with streaming or interaction-oriented capabilities and evaluates them across backend runtime, media transport, client-observed latency, and interaction behavior. With suitable backend selection and the WebRTC path, the tested system reaches approximately 0.9 to 1.0 s to first VLM text and 1.3 to 1.5 s to first non-silent TTS audio, while exposing backend adaptation costs and differences in real-time interaction behavior.
Comments12 pages. Submitted to IBC 2026