arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SNF-Bench:分离长视野固定相机视频生成中的静态漂移与自然流动

SNF-Bench: Separating Static Drift from Natural Flow in Long-Horizon Fixed-Camera Video Generation

Matiur Rahman Minar, Seunghun Oh, Ganghyeon Jeong, Unsang Park

arXiv 2608.28694首次发表:更新:

发表机构

Sogang University(西江大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

SNF-Bench 是针对长视野固定相机视频生成的评估框架,可分离静态漂移与自然流动,实验发现全帧运动与静态区域漂移对输出排序几乎相反,其指标能针对性评估对应因素。

AI 中文摘要

长视野视频生成采用全帧指标评估,这些指标会奖励运动和时间一致性。对于固定相机拍摄的自然场景,这会产生一种歧义:水、火、烟或雨的运动是理想的,而背景的运动则是错误的。因此,一个系统可能在运动方面得分很高,但其场景却发生漂移;或者在一致性方面得分很高,但其流动却停滞不前。我们推出 SNF-Bench,这是一个用于长视野固定相机生成的评估框架,它将每个场景划分为静态支撑和动态流动,并分别报告静态保真度、具有绝对幅度的流动持久性以及漂移泄漏,从不将其作为一个单一分数。漂移泄漏是具有解释性的上下文,而非标题测量值。每个因素都通过机械方式进行验证,而非通过与偏好的相关性:我们将已知严重程度的全局平移、旋转和缩放漂移以及渐进式后期冻结注入真实生成结果中,并要求每个因素按其规定方向响应,且对其未针对的损坏保持选择性。在一种记录的常见推理配置下审计公开发布的长视野文本条件检查点,加上带有发布管道参考的图像条件轨迹和部署敏感性面板,我们发现全帧运动和静态区域漂移对相同输出产生了几乎相反的排序。在最大受控平移下,fBD 和 NBF 上升至基线的 1.86 倍和 1.32 倍,但全帧动态度仅达到 1.07 倍——奖励了这种损坏。SNF-Bench 测量运动发生的位置以及它是否持续存在;它不测量物理真实性。项目页面:this https URL。

英文摘要

Long-horizon video generation is evaluated with whole-frame metrics that reward motion and temporal consistency. For fixed-camera nature scenes this creates an ambiguity: motion of water, fire, smoke, or rain is desirable, whereas motion of the background is an error. A system can therefore score well on motion while its scene drifts, or on consistency while its flow stagnates. We introduce SNF-Bench, an evaluation framework for long-horizon fixed-camera generation that partitions each scene into static support and dynamic flow and reports static fidelity, flow persistence with absolute magnitude, and drift leakage separately, never as one score. Drift leakage is interpretive context rather than a headline measurement. Each factor is validated mechanistically rather than by correlation with preference: we inject global translation, rotation, and scale drift and progressive late freezing at known severity into real generations, and require each factor to respond in its stated direction and to remain selective against corruptions it does not target. Auditing publicly released long-horizon text-conditioned checkpoints under one recorded common inference configuration, plus an image-conditioned track with released-pipeline references and a deployment-sensitivity panel, we find that whole-frame motion and static-region drift induce near-opposite orderings of the same outputs. At maximum controlled translation, fBD and NBF rise to $1.86\times$ and $1.32\times$ baseline, but whole-frame Dynamic Degree reaches only $1.07\times$---rewarding the corruption. SNF-Bench measures where motion occurs and whether it persists; it does not measure physical realism. Project page: https://minar09.github.io/snfbench/.

CommentsProject page: https://minar09.github.io/snfbench/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑