arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

PolarScale:面向辐射度一致RGB到Stokes估计的物理基准

PolarScale: A Physics-Grounded Benchmark for Radiometrically Consistent RGB-to-Stokes Estimation

Beibei Lin, Tingting Chen, Xin Zhang, Wenhao Zhao, Dongjun Li, Zifeng Yuan

arXiv 2610.08346首次发表:更新:

发表机构

National University of Singapore; University of Michigan, Ann Arbor(新加坡国立大学; 密歇根大学安娜堡分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

PolarScale提出物理基准,将RGB到Stokes估计中的辐射尺度作为显式预测目标,通过多模型评估表明恢复模型优于生成模型,并提升下游视觉任务性能。

AI 中文摘要

偏振成像提供了超越强度成像的物理线索,但通常需要专用硬件。近期方法从RGB类输入推断偏振信息,却仅预测归一化的Stokes分量或相对描述符,而完整Stokes重建所需的辐射度量已被从中消去。我们提出PolarScale基准,将该尺度作为明确的预测与评估目标。基于现有的三色全Stokes测量,PolarScale采用每场景归一化总强度图像$s_0$(场景参考线性图像,而非消费级sRGB照片),要求模型预测归一化Stokes分量、AoLP/DoLP/DoCP以及每场景尺度。由于尺度已从输入中消去,其在物理上不可辨识;因此PolarScale将数据集条件下的语义尺度估计与恒定尺度对照进行比较,并辅以角度、自一致性和物理边界指标。在七个基于恢复和生成的骨干网络及三种预测策略中,最强的恢复模型以3.6-4.3%的平均相对误差估计尺度(恒定对照为5.7%),且违反物理边界的像素少于0.25%,而两个生成基线则坍缩至接近零的尺度;显式描述符监督提升了描述符精度(MAE的PSNR从18.88 dB提升至23.66 dB)。预测的全Stokes表示改善了漫反射/镜面反射分离、材质分割和眩光分类,尽管在漫反射/镜面反射分离中,学习到的尺度仅与恒定对照表现相当。

英文摘要

Polarization imaging provides physical cues beyond intensity imaging but typically requires specialized hardware. Recent methods infer polarization from RGB-like inputs, yet predict only normalized Stokes components or relative descriptors, from which the radiometric scale needed for full Stokes reconstruction has been divided out. We introduce PolarScale, a benchmark that makes this scale an explicit prediction and evaluation target. Built on existing trichromatic full-Stokes measurements, PolarScale takes the per-scene normalized total-intensity image $s_0$ (a scene-referred linear image, not a consumer sRGB photograph) and asks models to predict normalized Stokes components, AoLP/DoLP/DoCP, and a per-scene scale. Because the scale is divided out of the input, it is not physically identifiable; PolarScale therefore evaluates dataset-conditioned semantic scale estimation against a constant-scale control, together with angular, self-consistency, and physical-bound metrics. Across seven restoration-based and generative backbones and three prediction strategies, the strongest restoration models estimate the scale with 3.6-4.3% mean relative error versus 5.7% for the constant control and violate physical bounds on fewer than 0.25% of pixels, whereas two generative baselines collapse to a near-zero scale; explicit descriptor supervision improves descriptor accuracy (23.66 vs. 18.88 dB PSNR for MAE). Predicted full-Stokes representations improve diffuse/specular separation, material segmentation, and glare classification, although in diffuse/specular separation the learned scale performs only on par with the constant control.

Comments22 pages, 17 figures, 8 tables. Accepted to NeurIPS 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑