arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

鲁棒多模态动态目标分割

Robust Multimodal Dynamic Object Segmentation

Zhe Xin, Hanzhi Chang, Penghui Huang, Yinian Mao, Guoquan Huang

arXiv 2607.18153首次发表:更新:

发表机构

Meituan UAV; University of Delaware(美团无人机; 特拉华大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对动态目标分割难题,提出整合多模态线索的框架,设计结合Transformer与特征聚类模块的网络进行分类,引入新后处理方法,在动态目标分割和静态场景重建任务中取得最优性能。

AI 中文摘要

动态目标分割在许多视觉应用中起着关键作用,如从动态视频进行静态场景重建。现有的基于光流的方法无法确保沿物体边界的静态/动态分割一致,基于3D重建的方法对重建误差高度敏感。为解决这些局限,我们提出一个动态目标分割框架,通过整合多模态线索生成精确完整的动态掩码。设计了结合Transformer架构与特征聚类聚合模块的网络进行多模态特征轨迹的静态/动态分类,还引入基于点查询的SAM后处理方法处理单个掩码内的多个物体。大量实验表明该方法在动态目标分割和静态场景重建任务中达到了当前最优性能。

英文摘要

Dynamic object segmentation plays a critical role in many visual applications such as static scene reconstruction from dynamic videos. However, existing optical flow-based methods fail to ensure consistent static/dynamic segmentation along object boundaries, while 3D reconstruction-based approaches are highly sensitive to reconstruction errors. To address these limitations, we present a dynamic object segmentation framework that can generate both precise and complete dynamic masks by integrating multimodal cues including 2D point tracks, 3D reconstruction, and semantic information. We design a network combining Transformer architectures with feature clustering aggregation modules to perform static/dynamic classification of multimodal feature trajectories. It enables the model to adaptively determine which type of feature should dominate based on the characteristics of each scene, while also mitigating the impact of feature degradation. Additionally, we introduce a novel point-query-based SAM post-processing method capable of handling multiple objects within a single mask. Extensive experiments demonstrate that our approach achieves state-of-the-art performance in both dynamic object segmentation and static scene reconstruction tasks.

CommentsAccepted by IEEE International Conference on Robotics & Automation ICRA 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑