arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于光场视差估计的频率结构场学习

Frequency-Structured Field Learning for Light-Field Disparity Estimation

Sara Monji-Azad, Yulin Liu, Jürgen Hesser

arXiv 2607.14941首次发表:更新:

发表机构

Mannheim Institute for Intelligent Systems in Medicine (MIISM), Medical Faculty Mannheim, Heidelberg University; Interdisciplinary Center for Scientific Computing (IWR), Heidelberg University; Central Institute for Computer Engineering (ZITI), Heidelberg University; CZS Heidelberg Center for Model-Based AI, Heidelberg University(曼海姆医学智能系统研究所(MIISM),曼海姆医学院,海德堡大学; 跨学科科学计算中心(IWR),海德堡大学; 计算机工程中央研究所(ZITI),海德堡大学; CZS海德堡基于模型的人工智能中心,海德堡大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究光场视差估计问题,提出FreqLF方法,通过EPI引导的傅里叶-局部框架,从全局和局部更新的潜在特征预测视差,避免显式构建视差体,实验表明该方法接近强监督基线精度,是有竞争力的替代方法。

AI 中文摘要

光场视差估计在平滑或无纹理区域需要全局一致性,在遮挡边界、细结构和深度突变处需要局部精度。现有方法通过EPI匹配、代价体或焦堆栈构建、视图聚合或直接卷积回归来满足这些要求,常依赖局部窗口、离散视差假设、内存密集型体或基于注意力的聚合。本文在场级制定视差估计,从全局和局部更新的EPI衍生潜在特征预测视差,无需显式构建视差体。引入FreqLF,一个EPI引导的傅里叶-局部框架,将水平和垂直EPI堆栈的角视差线索与中心视图外观特征编码在一起。这些线索投影到潜在场并通过堆叠的混合傅里叶-局部层更新。傅里叶低模更新实现全局特征交互,局部卷积保留精细视差细节所需的空间变化。坐标条件高斯混合解码器然后预测视差,使用混合均值作为最终估计。在HCI 4D光场基准上的实验表明,FreqLF接近强监督基线的精度,同时避免在基础模型中显式构建代价体。消融实验证实了傅里叶和局部分支的互补作用,缩放实验展示了跨空间分辨率的实际行为。这些结果表明,傅里叶-局部潜在场学习是光场视差估计的有竞争力的替代方法。代码即将发布。

英文摘要

Light-field disparity estimation requires global consistency in smooth or textureless regions and local precision near occlusion boundaries, thin structures, and abrupt depth transitions. Existing methods address these requirements through EPI matching, cost-volume or focal-stack construction, view aggregation, or direct convolutional regression, often relying on local windows, discrete disparity hypotheses, memory-intensive volumes, or attention-based aggregation. We instead formulate disparity estimation at the field level, predicting disparity from globally and locally updated EPI-derived latent features without explicitly constructing a disparity volume. We introduce FreqLF, an EPI-guided Fourier-local framework that encodes angular parallax cues from horizontal and vertical EPI stacks together with central-view appearance features. These cues are projected into a latent field and updated through stacked hybrid Fourier-local layers. Fourier low-mode updates enable global feature interaction, while local convolutions preserve spatial variations needed for fine disparity detail. A coordinate-conditioned Gaussian-mixture decoder then predicts disparity, using the mixture mean as the final estimate. Experiments on the HCI 4D Light Field Benchmark show that FreqLF approaches the accuracy of strong supervised baselines while avoiding explicit cost-volume construction in the base model. Ablations confirm the complementary roles of the Fourier and local branches, and scaling experiments demonstrate practical behavior across spatial resolutions. These results suggest that Fourier-local latent field learning is a competitive alternative for light-field disparity estimation. The code will be published soon.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑