arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.10321math.DG

不连续的先验模式部分与固有图像分解中模糊性的几何结构

Discontinuous Prior-Mode Sections and the Geometry of Ambiguity in Intrinsic Image Decomposition

Ziheng Chen, Liangchen Liu, Qishi Zhan, Minxuan Hu, Liang Geng

首次发表
浏览论文内容

中文总结 AI 辅助

研究固有图像分解中因模糊性产生的问题,提出基于先验模式部分在模糊性边界处不连续切换的几何解释,通过预测的可观察特征验证,在多个架构和数据集上得到相同特征,揭示了相关规律。

中文摘要 AI 辅助

2015年疯传的照片《那条裙子》因存在模糊性而使观察者分为两个阵营:相同的图像颜色既可以解释为在一种光照下的蓝黑色表面,也可以解释为在另一种光照下的白金表面。我们提出一种几何解释,即这种模糊性源于固有图像分解(将观察到的图像分离为反射率和光照的逆问题)中的奇异性。我们的核心主张是,先验模式部分(即先验偏好分解)在图像空间中的模糊性边界处切换,并且任何平滑的学习模型只能通过形成一个薄的过渡层来近似这种不连续的切换。这预测了两个可观察到的特征,其中$\Delta$是先验模式部分各分支之间的跳跃,$\lambda$是正则化强度:对于逆分解器,反照率雅可比矩阵按$|\Delta|/\sqrt{\lambda}$缩放;对于前向编码器,Fernet曲率在$1/\sqrt{\lambda}$的尺度上爆炸。在CGIntrinsics($N = 1998$张图像,$n = 2\times 10^7$像素)上,Careaga DPT的色温反照率雅可比矩阵与密集的地面真值反照率误差的部分斯皮尔曼相关性$r = 0.41$,而亮度和饱和度控制的相关性分别为$r = 0.087$和$r = 0.021$。在《那条裙子》上,CLIP ViT-L/14在$6473\,\mathrm{K}$处表现出潜在曲率峰值$\kappa = 73.03$,距离D65日光采样一步,而对照裙子图像在$\kappa = 34.75$处达到峰值且在D65附近没有可比特征。相同的特征出现在各种架构(U-Net逆变器、扩散逆变器、ViT编码器)和数据集(渲染的室内场景、网络照片)中,每个都用适合其模型类别的可观察量进行测量。

英文摘要

The viral 2015 photograph known as "The Dress" divides observers into two camps because it is ambiguous: the same image colors can be explained either as a blue-black surface under one illuminant or as a white-gold surface under another. We propose a geometric account in which the ambiguity arises from a singularity in intrinsic image decomposition, the inverse problem of separating an observed image into reflectance and illumination. Our central claim is that the prior-mode section, i.e. the prior-preferred decomposition, switches across an ambiguity boundary in image space, and that any smooth learned model can only approximate this discontinuous switch by forming a thin transition layer. This predicts two observable signatures, where $Δ$ is the jump between branches of the prior-mode section and $λ$ is the regularization strength: for inverse decomposers, an albedo Jacobian scaling as $|Δ|/\sqrtλ$; and for forward encoders, the Fernet curvature that blows up on a scale of $1/\sqrtλ$. On CGIntrinsics ($N=1998$ images, $n=2\times 10^7$ pixels), the color-temperature albedo Jacobian of Careaga DPT has partial Spearman correlation $r=0.41$ with dense ground-truth albedo error, compared with $r=0.087$ and $r=0.021$ for brightness and saturation controls. On "The Dress", CLIP ViT-L/14 exhibits a latent curvature peak of $κ=73.03$ at $6473\,\mathrm{K}$, one sampled step from D65 daylight, while a control dress image peaks at $κ=34.75$ with no comparable feature near D65. The same characteristic appears across architectures (U-Net inverter, diffusion inverter, ViT encoder) and datasets (rendered indoor scenes, web photograph), each measured with the observable appropriate to its model class.

↑