arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

从不相交数据进行统一视频密集预测

Unified Video Dense Prediction from Disjoint Data

Yihong Sun, Seoung Wug Oh, Jiahui Huang, Bharath Hariharan, Joon-Young Lee

arXiv 2607.21592首次发表:更新:

发表机构

Adobe Research; Cornell University(Adobe 研究院; 康奈尔大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究旨在解决现有任务特定注释分散问题,提出统一视频模型UniD,从不相交特定领域数据集联合预测八个密集场景属性。通过简单蒸馏步骤,利用预训练扩散模型视觉先验弥合领域差距,性能优于特定任务专家和多任务基线,泛化能力强。

AI 中文摘要

场景理解需要同时预测几何、外观和语义。然而,现有的特定任务注释分散在不兼容的特定领域数据集中。当前的统一系统通过将训练限制在完全共同注释的数据上,或通过产生伪标签的巨大计算成本来规避这一问题。为了缓解这一问题,我们引入了UniD,一个统一的视频模型,它联合预测八个密集场景属性——深度、表面法线、语义分割、边界、人体部位、反照率、阴影和材质——所有这些都从不相交的特定领域数据集中学习。我们提出了一个简单而有效的蒸馏步骤,其中每个任务的专家通过轻量级任务投影仪监督统一的主干,消除了对注释重叠或伪标签的需求。我们的关键见解是,预训练扩散模型的强大视觉先验足以弥合不相交训练源引入的领域差距,从而能够对训练期间从未见过的场景任务组合进行稳健泛化。UniD在与特定任务专家和多任务基线的比较中取得了有竞争力的性能,对分布外场景具有强大的泛化能力,并增强了时间和跨任务的一致性。代码和视频结果可在这个https URL上获取。

英文摘要

Scene understanding requires simultaneous prediction about geometry, appearance, and semantics. However, existing task-specific annotations are fragmented across incompatible, domain-specific datasets. Current unified systems circumvent this by restricting training to fully co-annotated data, or by incurring the large computational cost of pseudo-labeling. To mitigate this, we introduce UniD, a unified video model that jointly predicts eight dense scene properties-depth, surface normals, semantic segmentation, boundaries, human parts, albedo, shading, and materials-all learned from disjoint, domain-specific datasets. We propose a simple yet effective distillation step in which per-task experts supervise a unified backbone through lightweight task projectors, eliminating the need for annotation overlap or pseudo-labeling. Our key insight is that the strong visual priors of a pretrained diffusion model are sufficient to bridge the domain gaps introduced by disjoint training sources, enabling robust generalization to scene-task combinations never seen during training. UniD achieves competitive performance against per-task specialists and multi-task baselines, with strong generalization to out-of-distribution scenarios and enhanced temporal and cross-task consistency. Code and video results are available at https://unid-video.github.io/.

CommentsECCV 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑