arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于线性注意力的像素空间扩散实现高效高质量深度估计

Efficient and High-Quality Depth Estimation via Pixel-Space Diffusion with Linear Attention

Bingde Liu, Wu Ran, Jinglei Zhang, Huanhuan Yuan, Chao Ma

arXiv 2608.30129首次发表:更新:

发表机构

Shanghai Jiao Tong University; Zhiyuan College, Shanghai Jiao Tong University; MoE Key Lab of Artificial Intelligence, Institute of AI, Shanghai Jiao Tong University(上海交通大学; 上海交通大学致远学院; 上海交通大学人工智能研究院(教育部人工智能重点实验室))

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出基于线性注意力的像素空间生成框架Lapis,通过由粗到细的层级结构解决生成式深度估计的高计算成本问题,在多分辨率基准上实现SOTA精度,推理延迟大幅降低。

AI 中文摘要

本研究提出了Lapis,这是一种基于线性注意力的像素空间生成框架,可通过单步扩散实现高效且高保真的深度估计。尽管生成框架已凭借卓越的细节保真度显著推进了单目深度估计,但标准注意力的O(N²)复杂度以及多步去噪过程在将其扩展至高分辨率图像应用时会带来难以承受的计算成本。尽管线性注意力与单步预测直观上可行,但直接应用会导致结构一致性差、细节丢失和噪声问题。Lapis通过由粗到细的层级结构纠正了这些局限:具体而言,补丁级一致性模块通过整合语义与空间先验恢复结构连贯性;随后,像素级细化模块通过基于跳跃连接的像素对应关系恢复清晰的几何边界;此外,为缓解单步扩散固有的采样噪声,本研究利用流形假设并采用直接x预测策略以适配干净数据流形。在多个基准上的广泛评估表明,Lapis在不同分辨率下均始终达到最先进(SOTA)的精度与边界清晰度,与此前最先进的生成模型相比,其在1080P分辨率下推理延迟最多降低7.6倍,在1440P分辨率下最多降低10.9倍。

英文摘要

This work presents $\textbf{Lapis}$, a $\textbf{l}$inear-$\textbf{a}$ttention-based $\textbf{pi}$xel-$\textbf{s}$pace generative framework that achieves efficient and high-fidelity depth estimation with one-step diffusion. While generative frameworks have significantly advanced monocular depth estimation with superior detail fidelity, the $\mathcal{O}(N^2)$ complexity of standard attention and the multi-step denoising process introduce prohibitive computational costs when scaling them to high-resolution image applications. Although linear attention and one-step prediction are intuitively viable, directly applying them leads to poor structural consistency, detail loss, and noise. Lapis rectifies these limitations through a coarse-to-fine hierarchy. Specifically, a Patch-level Consistency Module restores structural coherence by integrating semantic and spatial priors. Subsequently, a Pixel-level Refinement Module recovers sharp geometric boundaries via skip-connection-based pixel correspondence. Furthermore, to mitigate sampling noise inherent in one-step diffusion, we leverage the manifold assumption and adopt a direct $\mathbf{x}$-prediction strategy to target the clean data manifold. Extensive evaluations on multiple benchmarks demonstrate that Lapis consistently achieves state-of-the-art (SOTA) accuracy and boundary sharpness across various resolutions, reducing inference latency by up to 7.6$\times$ at 1080P and 10.9$\times$ at 1440P resolution compared to previous SOTA generative models.

CommentsAccepted by ECCV 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑