arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Hyper-RED:通过语义超图蒸馏实现可扩展的事件预训练

Hyper-RED: Scalable Event Pre-training via Semantic Hypergraph Distillation

Meisen Wang, Zhiqiang Tian, Wei Bao, Chengjie Wang, Shaoyi Du, Siqi Li

arXiv 2609.16811首次发表:更新:

发表机构

Tsinghua University; Inner Mongolia Agricultural University(清华大学; 内蒙古农业大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对事件数据稀缺和模态差异问题,提出Hyper-RED框架,利用超图建模高阶语义关联并蒸馏图像知识至事件编码器,在五个数据集上实现可扩展且最先进的性能。

AI 中文摘要

事件相机在鲁棒视觉感知方面展现出巨大潜力,但由于大规模标注事件数据的稀缺,扩展事件表示学习仍然具有挑战性。预训练图像模型提供了可扩展的语义监督,但现有的图像到事件方法依赖于刚性的像素级或令牌级对齐,忽视了纹理、密度和外观方面的模态差异,可能导致语义崩溃并限制可迁移性。为解决这一问题,我们提出了Hyper-RED,一个简单、无痛且可扩展的图像到事件预训练框架,将高阶语义结构从图像迁移到事件。Hyper-RED使用超图来建模和对齐多个图像与事件令牌之间的高阶语义关联,实现跨模态知识迁移,同时适应模态特定差异,而非强制刚性的一一对应。具体而言,给定一个配对的事件-图像样本,Hyper-RED利用DINOv3提取空间令牌表示,并构建图像、事件和跨模态语义超图,其中每条超边连接多个语义相关的令牌。我们进一步引入超图关系蒸馏损失,施加互补的模态内和跨模态约束,使事件编码器能够继承源自图像的语义组织,同时保留局部关系一致性和事件特定特征。在五个事件数据集上的三项任务实验中,从ViT-S到ViT-L均展示出一致的可扩展性和最先进的性能(图1)。代码可在以下网址获取:此https URL。

英文摘要

Event cameras have shown great potential for robust visual perception, yet scaling event representation learning remains challenging due to the scarcity of large-scale annotated event data. Pretrained image models provide scalable semantic supervision, but existing image-to-event methods rely on rigid pixel-wise or token-wise alignment that overlooks modality discrepancies in texture, density, and appearance, potentially causing semantic collapse and limiting transferability. To address this issue, we propose Hyper-RED, a simple, painless, and scalable image-to-event pretraining framework that transfers high-order semantic structures from images to events. Hyper-RED uses hypergraphs to model and align high-order semantic associations among multiple image and event tokens, enabling cross-modal knowledge transfer while accommodating modality-specific differences rather than enforcing rigid one-to-one correspondence. Specifically, given a paired event--image sample, Hyper-RED leverages DINOv3 to extract spatial token representations and constructs image, event, and cross-modal semantic hypergraphs, where each hyperedge connects multiple semantically correlated tokens. We further introduce a hypergraph relational distillation loss that imposes complementary intra- and cross-modal constraints, enabling the event encoder to inherit image-derived semantic organization while preserving local relational consistency and event-specific characteristics. Experiments on three tasks across five event datasets demonstrate consistent scaling from ViT-S to ViT-L and state-of-the-art performance (Fig.1). The code is available at: https://github.com/meisenwang/Hyper--RED.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑