发表机构
College of Artificial Intelligence, Zhejiang University; Zhongguancun Academy(浙江大学人工智能学院; 中关村学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
TransNormal-2提出基于FLUX.2的修正流框架,通过几何感知损失和边缘感知解码校正VAE重建退化,实现精确单目法线估计,在通用基准上超越MoGe-2且仅需少量标注。
AI 中文摘要
基于扩散的模型能够实现单目几何估计,但其像素级精度受到一个共同且研究不足的误差源限制:VAE重建退化。VAE编码器-解码器中的8倍空间压缩会降低物体边界处的表面法线质量;即使对真实法线进行编码和解码,也会引入1.3–8.5°的平均角度误差(MAE),其中边缘MAE达到全局MAE的2.8倍。我们提出了TransNormal-2,一个基于FLUX.2的修正流框架,采用单步确定性推理,从VAE解码器的两侧解决这一退化问题:在训练期间如何监督潜在预测,以及在推理时如何校正解码后的法线。首先,几何感知的像素空间损失,包括逆渲染自一致性、von Mises-Fisher角度损失和小波边缘感知正则化,通过强制球面法线几何和漫反射图像形成线索(在VAE解码后)来补充潜在MSE。其次,一个轻量级的几何精化模块(GRM)应用RGB引导的残差校正,以减少边界局部的解码误差,而不会随意重写粗略预测。在通用场景基准上,TransNormal-2在所有八项报告指标上匹配或超过MoGe-2,同时仅使用其1.4%的任务特定法线标注。对于透明物体,增益最为明显,在ClearGrasp上将MAE降低了4.2°,在ClearPose上比最强先前基线降低了3.1°。代码将在此https URL发布。
英文摘要
Diffusion-based models enable monocular geometry estimation, yet their pixel-space precision is limited by a shared, under-studied error source: VAE reconstruction degradation. The 8x spatial compression in the VAE encoder-decoder degrades surface normals at object boundaries; even encoding and decoding ground-truth normals introduces 1.3--8.5° of mean angular error (MAE), with edge MAE reaching 2.8x the global MAE. We present TransNormal-2, a FLUX.2-based rectified-flow framework with single-step deterministic inference that addresses this degradation on both sides of the VAE decoder: in how latent predictions are supervised during training, and in how decoded normals are corrected at inference. First, geometry-aware pixel-space losses, including inverse rendering self-consistency, von~Mises-Fisher angular loss, and wavelet edge-aware regularization, complement latent MSE by enforcing spherical normal geometry and diffuse image-formation cues after VAE decoding. Second, a lightweight Geometric Refinement Module (GRM) applies an RGB-guided residual correction to reduce boundary-localized decoding errors without freely rewriting the coarse prediction. On general-scene benchmarks, TransNormal-2 matches or exceeds MoGe-2 on all eight reported metrics while using only 1.4% as many task-specific normal annotations. The gains are clearest for transparent objects, reducing MAE by 4.2° on ClearGrasp and 3.1° on ClearPose over the strongest prior baselines. Code will be released at https://longxiang-ai.github.io/TransNormal-2.
CommentsProject Page: https://longxiang-ai.github.io/TransNormal-2