arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.38476cs.CV

为任务特定视觉感知策划合成数据

Curating Synthetic Data for Task-Specific Visual Perception

Saptarshi Neil Sinha, Paul Julius Kühn, Michael Weinmann

首次发表
浏览论文内容

中文总结 AI 辅助

本文探讨为专用视觉系统策划合成数据,提出程序化渲染、物理模拟和生成式AI三种范式,并论证混合流程是实现可靠模拟到现实迁移的关键。

中文摘要 AI 辅助

合成数据在通用数据集无法提供任务所需的领域特定先验知识,且人工标注昂贵、不精确或不可行的情况下最具价值。在本文中,我们认为专用视觉系统的核心问题不在于如何生成更多数据,而在于应生成哪些数据。因此,我们讨论了策划合成数据,其场景内容、外观变化、传感特性和标注均围绕给定任务精心设计。我们考察了三种互补的策划范式。程序化渲染提供对场景参数的显式控制,标注随之由构造产生。基于物理的模拟编码了观察效应背后的机制,并产生精确对齐的监督对。生成式AI从少量真实种子集学习传感器特定外观,达到合理的逼真度,但仍易产生幻觉和不准确标注。这些范式通过工业表面缺陷检测、退化数字化奥托克罗姆胶片修复以及基于RGB和事件数据的6DoF姿态估计等示例加以说明。利用这些示例,我们分析了跨这些范式结合合成与真实数据的不同数据体制和训练策略。我们得出结论,策划合成数据最好被理解为对真实观测的补充,而结合可控监督与学习外观的混合流程是实现可靠模拟到现实迁移的最有前景方向。

英文摘要

Synthetic data are most valuable where general-purpose datasets cannot provide the domain-specific priors a task requires, and where manual annotation is expensive, imprecise, or infeasible. In this article we argue that the central question for specialized vision systems is not how to generate more data, but which data to generate. We therefore discuss curated synthetic data, whose scene content, appearance variations, sensing characteristics, and annotations are deliberately designed around a given task. We examine three complementary curation paradigms. Procedural rendering offers explicit control over scene parameters and the annotations follow by construction. Physically-based simulation encodes the mechanism behind an observed effect and yields exactly aligned supervision pairs. Generative AI learns sensor-specific appearance from small real seed sets and attains plausible realism, though it remains prone to hallucination and to inaccurate annotation. These paradigms are illustrated with examples from industrial surface defect detection, restoration of degraded digitized autochrome plates, and 6DoF pose estimation from RGB and event data. Using these examples, we analyze different data regimes and training strategies that combine synthetic and real data across these paradigms. We conclude that curated synthetic data are best understood as a complement to real observations, and that hybrid pipelines combining controllable supervision with learned appearance are the most promising direction for reliable sim-to-real transfer.

发表机构

  • Fraunhofer IGD(弗劳恩霍夫计算机图形研究所)
  • Delft University of Technology(代尔夫特理工大学)

机构由 AI 辅助整理,请以论文原文为准。

↑