发表机构
ETH Zurich(苏黎世联邦理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对现有文档展开数据集规模和质量不足的问题,本文提出大规模合成数据集SyntheticDoc,含百万级高分辨率样本及精确标注,通过物理模拟和路径追踪生成,实验证明其能提升基线模型性能。
AI 中文摘要
深度学习模型已成为文档校正和光照校正的标准工具,然而其性能从根本上受限于训练数据。近十年来,社区严重依赖Doc3D,这是一个开创性但规模和质量日益受限的文档展开数据集。为解决这一瓶颈,我们引入了SyntheticDoc,一个大规模、高质量的数据集,旨在推动文档展开的边界。SyntheticDoc由1,000,000个高分辨率程序化生成的训练样本以及广泛的验证集和测试集组成。每个样本都配有丰富的、像素级精确的标注,包括UV图、法线图、反照率和着色。为确保物理准确性和照片级真实感,纸张几何体通过基于物理的模拟器生成,并使用路径追踪器渲染。为展示我们数据集的优势,我们在SyntheticDoc上训练了一个简单的基线模型,并报告其在文档展开和光照校正任务上与最先进方法的性能比较。我们的数据集可在以下网址获取:https URL,生成代码可在以下网址获取:https URL。
英文摘要
Deep learning models have become the standard tool for document rectification and illumination correction, yet their performance is fundamentally bound by their training data. For nearly a decade, the community has heavily relied on Doc3D, a pioneering but increasingly limited document unwarping dataset in terms of scale and quality. To address this bottleneck, we introduce SyntheticDoc, a massive, high-quality dataset designed to push the boundaries of document unwarping. SyntheticDoc is composed of 1,000,000 high-resolution procedurally generated training samples, alongside extensive validation and test sets. Each sample is paired with rich, pixel-perfect annotations, including UV maps, normal maps, albedo and shading. To ensure physical accuracy and photorealism, the paper geometries are generated via a physics-based simulator and rendered using a path tracer. To demonstrate the benefit of our dataset, we train a simple baseline model on SyntheticDoc and report on its performance in comparison to state-of-the-art methods on both document unwarping and illumination correction tasks. Our dataset is available at https://igl.ethz.ch/projects/SyntheticDoc/ and the code used to generate it at https://github.com/tanguymagne/SyntheticDoc .
CommentsD. Woortmann and T. Magne -- Equal contribution. Accepted at ECCV 2026 (Spotlight). 20 pages
Journal refLecture Notes in Computer Science, vol. 17057, pp. 131-150, Springer (2026)
DOI:10.1007/978-3-032-37335-9_8