arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

PhysMLLMs:用于图像与视频的统一指称分割及接地推理的空间先验

PhysMLLMs: Spatial Priors for Unified Referring Segmentation and Grounded Reasoning of Images and Videos

Siyao Yan, Bo Han, Jisheng Dang, Bimei Wang, Shude Wang, Hong Peng, Yulan Guo, Jianhuang Lai, Bin Hu, Tat-SengChua

arXiv 2608.24574首次发表:更新:

发表机构

School of Information Science and Engineering, Lanzhou University; School of Electronics and Communication Engineering, Sun Yat-sen University; School of Computing, National University of Singapore(兰州大学信息科学与工程学院; 中山大学电子与通信工程学院; 新加坡国立大学计算机学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

PhysMLLMs是注入物理启发空间先验的训练架构,通过REPA-Global机制提升视频分割的时空一致性,且不损害图像级接地与通用多模态能力。

AI 中文摘要

视频多模态大语言模型支持语言引导的视频分割,但常存在时空不一致问题,如抖动、漂移和身份切换。当目标部分被遮挡或出现相似物体时,这类失效更常见,原因可能是当前训练缺乏显式空间先验,导致难以在时间维度上维持稳定的空间身份与形状。我们提出PhysMLLMs,一种训练阶段的先验注入架构,将受物理启发的空间连续性先验注入视频多模态大语言模型。PhysMLLMs旨在通过训练期间将学生全局视觉表示与冻结的教师模型对齐,鼓励更稳定的以对象为中心的表示。其核心机制为全局表示先验对齐(REPA-Global),利用离线嵌入缓存和调度蒸馏计划,从冻结的DINOv2教师模型中蒸馏全局视觉表示。该设计保持推理阶段不变,不增加推理时间成本。在多个视频基准测试中,PhysMLLMs提升了视频分割掩码质量和跨帧一致性,在涉及小目标、快速运动、遮挡、干扰项及推理查询的挑战性案例中增益更显著。在单帧指称图像分割和代表性通用VLM基准上,PhysMLLMs保持了可比性能,表明注入的空间先验可提升视频一致性,同时不损害图像级接地或通用多模态能力。代码可在指定网址获取。

英文摘要

Video multimodal large language models support language guided video segmentation, but they often show spatio temporal inconsistencies, e.g., jitter, drift, and identity switches. These failures are more common when targets are partly hidden or when similar objects appear nearby.One likely reason is that current training lacks explicit spatial priors, which makes it difficult to maintain stable spatial identity and shape over time. We present PhysMLLMs, a training-stage prior injection architecture that injects physics-inspired spatial continuity priors into Video MLLMs. PhysMLLMs is designed to encourage more stable object-centered representations by aligning the student global visual representation with a frozen teacher model during training. Our core mechanism, Global Representation Prior Alignment (REPA-Global), distills global visual representations from a frozen DINOv2 teacher using an offline embedding cache and a scheduled distillation plan. This design keeps inference unchanged and does not add inference time cost. Across multiple video benchmarks, PhysMLLMs improves video segmentation mask quality and cross-frame consistency, with larger gains on challenging cases involving small targets, fast motion, occlusion, distractors, and reasoning queries. On single-frame referring image segmentation and representative general VLM benchmarks, PhysMLLMs maintains comparable performance, demonstrating that the injected spatial prior improves video consistency without compromising image-level grounding or general multimodal capability. These results suggest that physics-inspired spatial prior injection can improve temporal stability while preserving general capability. The code is available at https://github.com/tusu-code/20260121-icml2026-2.git.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑