发表机构
Zhejiang University; The University of Hong Kong(浙江大学; 香港大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
E-WAVE提出无相关框架,利用全局注意力与贝塞尔曲线扭曲,实现事件流的高时间分辨率光流估计,降低轨迹误差25%。
AI 中文摘要
时间密集的光流对于沉浸式VR/AR系统中的动态感知至关重要,其中头部、手部和物体的快速运动必须被连续捕获和跟踪。现有的基于帧的光流估计方法受限于时间分辨率和计算成本之间的权衡;而事件相机凭借其高时间分辨率和能量效率,成为解决这一困境的自然方案。然而,基于事件的方法通常依赖相关体积来捕获成对体素对应关系,这会带来大量的内存和计算开销。我们提出了E-WAVE,一种用于从事件流进行高时间分辨率(HTR)光流估计的无相关框架。E-WAVE不构建全对相关体积,而是采用全局注意力机制来建模长距离特征依赖,并使用贝塞尔曲线进行轨迹引导的特征扭曲。通过迭代更新,它预测轨迹,从而允许在任意时间戳进行查询,而无需重复推理。在MultiFlow和DSEC-Flow上的实验表明,相对于最先进的基线,轨迹误差降低了25%,端点光流估计精度相当。使用头戴式原型在自采集数据上的额外评估验证了E-WAVE在具有挑战性的现实条件下保持鲁棒性。
英文摘要
Temporally dense optical flow is essential for dynamic perception in immersive VR/AR systems, where rapid head, hand, and object motion must be continuously captured and tracked. Existing frame-based optical flow estimation methods are constrained by the tradeoff between temporal resolution and computational cost; while event cameras, with their high temporal resolution and energy efficiency, serve as a natural solution to the dilemma. However, event-based approaches commonly rely on correlation volumes to capture pairwise voxel correspondences, which incur substantial memory and computation overhead. We present E-WAVE, a correlation-free framework for high-temporal-resolution (HTR) optical flow estimation from event streams. Instead of constructing all-pairs correlation volumes, E-WAVE employs global attention mechanism to model long-range feature dependencies and performs trajectory guided feature warping using Bézier curve. Through iterative updates, it predicts trajectories that allow for querying at arbitrary timestamps without repeated inference. Experiments on MultiFlow and DSEC-Flow demonstrate a 25% lower trajectory error and comparable endpoint flow estimation accuracy relative to state-of-the art baselines. Additional evaluations on self-captured data using a head-mounted prototype validate that E-WAVE remains robust under challenging real-world conditions.