arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

WHTMix:通过沃尔什-哈达玛令牌混合实现高效立体深度估计

WHTMix: Efficient Stereo Depth Estimation via Walsh-Hadamard Token Mixing

Prathyush Sajith, Emadeldeen Hamdan, Ahmet Enis Cetin

arXiv 2607.25234首次发表:更新:

发表机构

University of Illinois at Chicago(伊利诺伊大学芝加哥分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对立体深度估计在高分辨率和低延迟要求下的问题,提出用沃尔什-哈达玛令牌混合器替代立体变压器的联合自注意力阶段,并引入混合对数视差损失函数,实现减少计算量、延迟及降低远距离物体误差。

AI 中文摘要

对于驾驶、机器人技术和增强现实中的立体深度估计,必须在严格的延迟预算下以高分辨率运行。然而,在基于变压器的匹配器中,聚合场景上下文的全局自注意力随像素数量呈二次增长,并主导运行时。我们表明,立体变压器的联合自注意力阶段可以由数据独立的沃尔什-哈达玛令牌混合器代替,该混合器以对数线性成本在变换域中全局混合令牌,同时保留执行左右对应的依赖数据的交叉注意力。在合成驾驶数据上,混合器在端点误差上与注意力基线匹配,同时将模型计算减少2.46倍,单图像推理延迟减少2.65倍。复杂性分析表明,这种优势受序列长度与通道宽度之比的控制,这解释了为什么高分辨率立体匹配是特别有利的设置,以及为什么分类变压器不是。我们在非立体长序列基准上证实了这种令牌到通道的缩放。此外,我们引入了一种混合对数视差损失函数,旨在对与远距离物体对应的小视差像素进行加权。这种方法减少了远距离物体的误差,而不会产生任何额外的计算开销。

英文摘要

Stereo depth estimation for driving, robotics and augmented reality must run at high resolution under tight latency budgets, yet in transformer-based matchers the global self-attention that aggregates scene context grows quadratically with the number of pixels and comes to dominate runtime. We show that the joint self-attention stage of a stereo transformer, whose role is to spread context across both views, can be replaced by a data-independent Walsh-Hadamard token mixer that mixes tokens globally in the transform domain at log-linear cost, while the data-dependent cross-attention that performs left-right correspondence is retained. On synthetic driving data the mixer matches the attention baseline in end-point error while reducing model compute by a factor of 2.46 and single-image inference latency by a factor of 2.65. A complexity analysis shows the benefit is governed by the ratio of sequence length to channel width, which explains why high-resolution stereo matching is a particularly favorable setting and why classification transformers are not; we confirm this token-to-channel scaling on non-stereo long-sequence benchmarks. Furthermore, we introduce a hybrid log-disparity loss function designed to up-weight small-disparity pixels corresponding to long-range objects. This approach reduces the error on distant objects without incurring any additional computational overhead.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑