FreeFlow:一种用于光流估计的无偏置分层Transformer
FreeFlow: A Bias-free Hierarchical Transformer for Optical Flow Estimation
浏览论文内容
中文总结 AI 辅助
FreeFlow提出无流特定归纳偏置的分层Transformer,结合三种注意力变体,在多个光流基准上取得最先进结果并保持内存效率。
中文摘要 AI 辅助
光流方法通常依赖任务特定的归纳偏置,如相关体、特征扭曲和迭代细化等,以达到高精度。尽管这些偏置有效,但它们将模型限制在预定义的启发式规则中,可能限制其表达能力,并导致更复杂的流程和额外的计算成本。我们提出FreeFlow,一种构建时不使用任何流特定组件的分层Transformer,而是采用单一的前馈编码器-解码器。FreeFlow结合了三种注意力变体:用于局部处理的窗口注意力、用于跨窗口信息交换的移位窗口注意力,以及在降低分辨率下运行的全局注意力。由此产生的架构随模型容量自然扩展,使得从小型到大型变体都能获得一致的精度提升。尽管缺乏标准的归纳偏置,FreeFlow在主要基准上取得了最先进的结果,包括Sintel(Clean/Final上的EPE为0.68/1.48)、KITTI-2015(Fl-all为3.23%)和Spring(1px误差为3.192),同时在1080p推理时保持内存效率。
英文摘要
Optical flow methods typically rely on task-specific inductive biases, such as correlation volumes, feature warping, and iterative refinement, among others, to reach high accuracy. While effective, such biases constrain the model to predefined heuristics, which can limit its expressivity and lead to more complex pipelines and additional computational cost. We present FreeFlow, a hierarchical transformer built without any flow-specific components, using instead a single feed-forward encoder--decoder. FreeFlow combines three attention variants: window attention for local processing, shifted-window attention for cross-window information exchange, and a global attention operating at a reduced resolution. The resulting architecture scales naturally with model capacity, enabling a consistent accuracy gain from small to large variants. Despite the absence of standard inductive biases, FreeFlow achieves state-of-the-art results on major benchmarks, including Sintel (0.68/1.48 EPE on Clean/Final), KITTI-2015 (3.23 Fl-all), and Spring (3.192 1px), while remaining memory efficient at 1080p inference.
发表机构
- AI Center, Lomonosov MSU(罗蒙诺索夫莫斯科国立大学人工智能中心)
- Lomonosov Moscow State University(罗蒙诺索夫莫斯科国立大学)
- MSU Institute for Artificial Intelligence(莫斯科国立大学人工智能研究所)
机构由 AI 辅助整理,请以论文原文为准。