发表机构
Fulbright University Vietnam; University of California, Los Angeles(富布赖特越南大学; 加利福尼亚大学洛杉矶分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出流形外细化(OMR)方法,用冻结的V-JEPA 2.1世界模型引导视频生成器,在低额外成本下提升了视频的物理一致性与语义 adherence指标。
AI 中文摘要
现代视频生成器常无法正确处理物理动态:物体漂浮、轨迹违反重力、接触消失。标准去噪和流匹配目标虽能拟合视觉数据分布,但未明确惩罚这类物理违规。现有改进方法可提升物理一致性,但通常会大幅增加推理或训练成本:候选选择方法需生成并评分多个视频,基于梯度的世界模型引导需反复解码和重编码中间估计,生成器内部细化添加扰动和重去噪循环,而后训练则需要精心整理的数据和额外优化。我们提出流形外细化(Off-Manifold Refinement, OMR),一种推理时方法,直接将世界模型反馈注入单条采样轨迹。在计划的中间常微分方程(ODE)步骤中,我们用适配器空间V-JEPA 2.1的惊喜能量梯度增强生成器速度。这种外部修正可将潜在变量从未校正的采样轨迹移向冻结预测器判定为更具物理合理性的区域,之后生成器从校正状态继续渲染。小型训练的潜在变量到嵌入的适配器使梯度在推理时保持可处理,视频生成器和世界模型均保持冻结。在固定的400个提示的VideoPhy-2详细子集上,OMR将基础Wan2.2-T2V-A14B采样器的联合语义 adherence与物理常识指标从47.0%提升至52.0%(绝对提升5.0个百分点,相对提升10.6%);在单独固定的50个提示的效率子集上,其运行时为基础的1.71倍,远低于奖励/搜索替代方法的乘法成本。项目页面:this https URL。
英文摘要
Modern video generators routinely fail at physical dynamics: objects float, trajectories violate gravity, contacts vanish. Standard denoising and flow-matching objectives fit visual data distributions but do not explicitly penalize such physical violations. Existing remedies can improve physical consistency, but typically add substantial inference or training cost. Candidate-selection methods generate and score multiple videos, while gradient-based world-model guidance repeatedly decodes and re-encodes intermediate estimates. Generator-internal refinement adds perturbation and re-denoising loops, whereas post-training requires curated data and additional optimization. We propose Off-Manifold Refinement (OMR), an inference-time method that instead injects world-model feedback directly into a single sampling trajectory. During scheduled middle ODE steps, we augment the generator velocity with the gradient of an adapter-space V-JEPA 2.1 surprise energy. This external correction can move the latent away from the uncorrected sampling trajectory and toward regions ranked as more physically plausible by the frozen predictor, after which the generator continues rendering from the corrected state. A small trained latent-to-embedding adapter keeps the gradient tractable at inference, and both the video generator and the world model remain frozen. On our fixed 400-prompt VideoPhy-2 detailed subset, OMR lifts the joint Semantic-Adherence-and-Physical-Commonsense metric from 47.0% to 52.0% (+5.0pp absolute, +10.6% relative) over the base Wan2.2-T2V-A14B sampler. On a separate fixed 50-prompt efficiency subset, it requires $1.71 \times$ the base runtime rather than the multiplicative cost of reward/search alternatives. Project page: https://itruonghai.github.io/omr.
CommentsAccepted at BMVC 2026. Project page: https://itruonghai.github.io/omr