arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

多智能体视频预测:用于动态场景预测的自校正条件帧

Multi-Agent Video Prediction: Self-Correcting Conditional Frames for Dynamic Scene Forecasting

Qixin Zhang, Ajay Kumar, Zhi-Li Zhang

arXiv 2609.25302首次发表:更新:

发表机构

University of Minnesota(明尼苏达大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对远程驾驶中网络延迟导致新物体出现时预测失效的问题,提出多智能体视频预测框架,结合连续预测与掩码引导条件重调节,实现语义恢复并保持质量与效率。

AI 中文摘要

传输延迟显著降低了实时交互感知系统中的用户体验质量。在远程驾驶中,维持可靠的视觉反馈对于动态网络变化下的安全操作至关重要。尽管视频预测为补偿短期传输延迟和实现近似零延迟流提供了一种有前景的方法,但仅依赖预测的方法在高度动态场景中仍然脆弱,尤其是在上行链路中断期间出现新物体时。为应对这些挑战,我们提出了一种多智能体视频预测框架,该框架将连续的边缘侧视频预测与轻量级掩码引导的条件帧重调节相结合。该框架由三个角色专业化智能体组成:一个用于低延迟视觉连续性的连续预测智能体,一个用于检测新出现物体的车辆侧触发智能体,以及一个利用稀疏掩码引导修复预测器条件状态的条件重调节智能体。这种设计使得在不进行全帧重传的情况下,能够对外生场景变化进行语义恢复。我们通过在真实5G通信轨迹下的基准视频数据上进行大量实验,验证了所提出的框架。结果表明,我们的方法在保持感知质量和实际运行效率的同时,在网络引起的干扰下改善了对新物体的语义恢复。

英文摘要

Transmission latency significantly degrades user quality of experience in real-time interactive perception systems. In remote driving, maintaining reliable visual feedback is critical for safe operation under dynamic network variability. Although video prediction offers a promising approach to compensate for short-term transmission delays and approximate near-zero-latency streaming, prediction-only methods remain vulnerable in highly dynamic scenes, especially when newly emerged objects appear during uplink outages. To address these challenges, we propose a multi-agent video prediction framework that combines continuous edge-side video prediction with lightweight mask-guided conditional frame reconditioning. The framework consists of three role-specialized agents: a continuous prediction agent for low-latency visual continuity, a vehicle-side trigger agent for detecting newly appeared objects, and a conditional reconditioning agent that repairs the predictor conditioning state using sparse mask guidance. This design enables semantic recovery of exogenous scene changes without requiring full-frame retransmission. We validate the proposed framework through extensive experiments on benchmark video data under realistic 5G communication traces. Results show that our method improves semantic recovery of novel objects while preserving perceptual quality and practical runtime efficiency under network-induced disruptions.

CommentsPublished in the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), 2026

Journal refProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Workshops, 2026, pp. 1087-1096

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑