发表机构
Leibniz University Hannover(汉诺威莱布尼茨大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对现有基于事件相机的强化学习方法的缺陷,提出FLEET特征提取器,通过高效令牌化等技术实现与传感器分辨率解耦的端到端学习,在新基准上超越SOTA且鲁棒性更优。
AI 中文摘要
事件相机生成异步、高频的数据流,提供空间稀疏信息且延迟低于传统相机,从原理上看,这些特性应非常适合控制算法设计,但该领域的强化学习研究仍有限,因为现有方法未能充分利用传感器的优势:基于帧的方法将事件聚合为稀疏网格,这使得特征提取器的计算成本与传感器分辨率绑定,且模糊了时间信息;同时,现有生成基线依赖轨迹数据来预训练模型。我们提出FLEET(Feature Learning from Events via Efficient Tokenization,即通过高效令牌化从事件中学习特征),一种直接处理事件序列的特征提取器,利用随机傅里叶特征和交叉注意力,我们的架构将可变长度的事件流压缩为固定大小的潜在表示,这使特征提取器主干的推理成本与传感器分辨率解耦,支持无需辅助损失的端到端学习。我们在新的高吞吐量基准上验证FLEET,结果表明我们基于序列的方法超越了SOTA性能,且对观测频率变化表现出更优的鲁棒性。
英文摘要
Event cameras generate asynchronous, high-frequency data streams offering spatially sparse information at lower latency than traditional cameras. In principle, these properties should be ideal for the design of control policies. However, reinforcement learning research in this field remains limited as existing approaches fail to fully exploit the sensor's properties. CNN-based methods negate the sensors benefits by aggregating events into sparse grids. This couples compute cost to sensor resolution and blurs the temporal information. Meanwhile, existing generative baselines rely on the availability of trajectory data to pretrain the model. We propose FLEET (Feature Learning from Events via Efficient Tokenization), a feature extractor that processes event sequences directly. Leveraging random Fourier features and cross-attention, our architecture compresses variable streams into fixed-size latent representations. This decouples inference cost of the feature extractor's backbone from the sensor's resolution, enabling end-to-end learning without auxiliary losses. We validate FLEET on a new, high-throughput benchmark. The results demonstrate that our sequence-based approach surpasses SOTA performance and exhibits superior robustness to variations in observation frequencies.