发表机构
The University of Tokyo; Rikkyo University; RIKEN(东京大学; 立教大学; 理化学研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过事件域到RGB域的蒸馏,揭示了事件数据诱导的基于边缘的归纳偏差,包括颜色不变性、形状偏差和高频噪声鲁棒性,并证明其作为下游任务的有用先验。
AI 中文摘要
在ImageNet上训练的卷积神经网络已知表现出对局部高频纹理的强烈偏好,这种归纳偏差转化为在现实世界环境中对分布变化的脆弱鲁棒性。相比之下,事件相机仅记录场景亮度的变化,因此非常适合捕捉轮廓信息;然而,由于事件领域缺乏诊断基准,事件相机数据赋予视觉模型的归纳偏差仍未得到充分探索。在本工作中,我们利用从事件领域到RGB领域的知识蒸馏,以利用RGB领域中丰富的评估工具包,并系统地剖析这种归纳偏差。我们的实验表明,从事件领域的蒸馏在RGB领域中诱导出颜色不变性、形状偏差和对高频噪声的鲁棒性。我们确定其潜在机制为模型抑制对高频纹理的依赖,同时获得对基于边缘的对象形状的更强依赖。这一假设得到早期层中颜色和空间信息处理方式变化的支持,以及一种频谱权衡:对高频成分缺失的鲁棒性与对依赖频带污染和几何结构破坏的脆弱性共存。我们进一步表明,这种归纳偏差不同于现有的鲁棒化方法,并且作为多种下游任务的有用先验,其中形状和轮廓信息与其他线索共同发挥作用。代码可在以下https URL获取。
英文摘要
Convolutional neural networks trained on ImageNet are known to exhibit a strong preference for local high-frequency texture, an inductive bias that translates into fragile robustness against distribution shifts in real-world environments. Event cameras, in contrast, record only changes in scene brightness and are therefore well suited to capturing contour information; however, due to the absence of diagnostic benchmarks in the event domain, the inductive bias that event-camera data instills in vision models has remained underexplored. In this work, we use knowledge distillation from the event domain to the RGB domain so as to exploit the rich evaluation toolkit available in the RGB domain and systematically dissect this inductive bias. Our experiments show that distillation from the event domain induces, in the RGB domain, color invariance, shape bias, and robustness to high-frequency noise. We identify the underlying mechanism as the model suppressing its dependence on high-frequency texture while acquiring a stronger dependence on edge-based object shape. This hypothesis is supported by changes in how color and spatial information are processed at the early layers, together with a spectral trade-off in which robustness to the absence of high-frequency components coexists with vulnerability to contamination of the relied-upon frequency bands and to disruption of geometric structure. We further show that this inductive bias differs from existing robustification methods and that it functions as a useful prior for diverse downstream tasks in which shape and contour information contribute alongside other cues. The code is available at https://github.com/snskysk/event2rgb-distillation .
CommentsAccepted at NeurIPS 2026. All authors contributed equally. Code: https://github.com/snskysk/event2rgb-distillation