AI 中文总结
研究提出时间反转成像范式,通过TRACE-HEI数据集及多模态推理方法,从热、紫外和可见光谱中残余物理印记推断过去人类与环境交互,为时间反转成像奠定基础,开辟场景理解新方向。
AI 中文摘要
我们引入了时间反转成像,这是一种从渐弱的多模态痕迹推断场景中刚刚发生之事的新范式。我们的目标不是外推或插值视频帧,而是从热、紫外和可见光谱中可观察到的残余物理印记推断过去的人类与环境交互。为研究此问题,我们展示了TRACE-HEI,这是首个用于时间反转成像的概念验证数据集。为建立基准,我们提出一种多模态推理方法,该方法提取检测到的痕迹的结构化文本描述,并用于约束视觉语言引导的扩散模型以重建合理的过去帧。实验表明,从渐弱痕迹推断近期事件具有挑战性,但当互补模态减少解决方案的模糊性时是可行的。这项工作为时间反转成像定义了首个计算和实验基础,架起了视觉、物理和生成推理之间的桥梁,并为超越即时观察的场景理解开辟了新方向。
英文摘要
We introduce time-reversed imaging, a new paradigm that infers what just happened in a scene from fading multimodal traces. Instead of extrapolating or interpolating video frames, our goal is to infer past human-environment interactions from residual physical imprints observable in thermal, ultraviolet, and visible spectra. To study this problem, we present TRACE-HEI, the first proof-of-concept dataset for time-reversed imaging, containing synchronized tri-modal video sequences of actions such as sitting, touching, moving objects, and liquid spills, captured across diverse materials and recorded up to three minutes after contact. To establish the benchmark, we propose a multimodal inference approach that extracts structured textual descriptions of detected traces and uses them to constrain a vision-language-guided diffusion model for reconstructing plausible past frames. Experiments show that inferring recent events from fading traces is challenging but feasible when complementary modalities reduce solution ambiguity. This work defines the first computational and experimental foundation for time-reversed imaging, bridging vision, physics, and generative reasoning, and opening new directions for scene understanding beyond instantaneous observation.