发表机构
Shanghai Jiao Tong University; Beihang University; DISCOVER Robotics(上海交通大学; 北京航空航天大学; DISCOVER机器人公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出Engram-E2VID框架,通过外观印迹生成激活实现参考式事件到视频重建,在三个基准测试中PSNR最高提升3.29 dB、LPIPS最高降低0.08,性能随重建间隔增加下降更慢。
AI 中文摘要
参考式事件到视频重建旨在从参考帧及参考帧到目标帧时间间隔内捕获的事件流中恢复目标RGB帧。尽管事件提供了细粒度的时间线索,但它们编码的是稀疏异步的对数强度变化而非绝对外观,因此忠实重建存在固有挑战。核心挑战在于将事件衍生的目标时刻结构与参考帧的相关外观信息关联,尤其在复杂运动和长时间间隔下。本研究提出Engram-E2VID,一种通过外观印迹的生成激活来重建目标帧的结构引导框架。具体而言,参考帧被编码为令牌空间外观印迹,而事件流和参考上下文被转换为捕获运动边界及事件诱导结构变化的目标时刻运动结构支架。在单步扩散骨干网络中,支架衍生的结构令牌逐层与相关外观印迹交互并激活它们。这种令牌空间关联使目标结构能够在不依赖直接像素级对应关系的情况下访问参考外观,同时扩散先验补充不确定或新显现的区域。在三个基准测试中,Engram-E2VID相较于最强的相同输入基线,PSNR提升最高达3.29 dB,LPIPS降低最高达0.08,且随重建间隔增加性能下降更缓慢。
英文摘要
Reference-based event-to-video reconstruction aims to recover target RGB frames from a reference frame and the event stream captured over the reference-to-target interval. Although events provide fine-grained temporal cues, they encode sparse and asynchronous log-intensity changes rather than absolute appearance, making faithful reconstruction intrinsically challenging. The central challenge lies in associating event-derived target-time structures with relevant appearance information from the reference frame, especially under complex motion and long temporal intervals. In this work, we propose Engram-E2VID, a structure-guided framework that reconstructs target frames through the generative activation of appearance engrams. Specifically, the reference frame is encoded into token-space appearance engrams, while the event stream and reference context are transformed into a target-time motion-structure scaffold that captures motion boundaries and event-induced structural changes. Within a one-step diffusion backbone, scaffold-derived structural tokens progressively interact with and activate relevant appearance engrams across layers. This token-space association allows target structures to access reference appearance without relying on direct pixel-wise correspondence, while the diffusion prior complements uncertain or newly revealed regions. Across three benchmarks, Engram-E2VID improves PSNR by up to 3.29 dB and reduces LPIPS by up to 0.08 over the strongest same-input baseline, while degrading more slowly as the reconstruction interval increases.
Comments9 pages, 5 figures