拉格朗日-哈密顿流用于视频预测和图像生成:辛几何视角
Lagrangian--Hamiltonian Flows for Video Prediction and Image Generation: A Symplectic Perspective
- Boston University(波士顿大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
提出LHFM几何框架,基于辛几何和哈密顿流建模图像动力学,用于视频预测和图像生成,在精度相近时实现最低FLOP计数,并优于流匹配基线。
AI中文摘要:
我们提出LHFM,一种用于学习图像动力学的几何框架。借鉴经典力学、辛几何和几何量子化中的核心结构,LHFM将每幅图像表示为精确的拉格朗日图,并通过依赖于图像的哈密顿流来建模其演化,从而得到图像速度的传输-源参数化。我们的主要应用是确定性视频预测:LHFM-V是一种循环模型,通过积分预测的传输场和源场来推进帧,在预测精度相近的比较循环模型中实现了最低的FLOP计数。图像变体LHFM-I表明,相同的构造与流匹配兼容:在匹配的实验中,它获得了比流匹配基线更低的FID。
英文摘要:
We introduce LHFM, a geometric framework for learning image dynamics. Drawing on structures central to classical mechanics, symplectic geometry, and geometric quantization, LHFM represents each image as an exact Lagrangian graph and models its evolution through image-dependent Hamiltonian flows, which yield a transport--source parameterization of image velocities. Our primary application is deterministic video prediction: LHFM-V is a recurrent model that advances frames by integrating predicted transport and source fields, and achieves the lowest reported FLOP count among the compared recurrent models with similar prediction accuracy. The image variant, LHFM-I, shows that the same construction is compatible with flow matching: in a matched experiment, it attains a lower FID than the flow-matching baseline.