关于光流模型的现实世界通用性
On the Real-World Generalisability of Optical Flow Models
浏览论文内容
中文总结 AI 辅助
研究光流模型在现实世界的通用性问题,构建现实世界评估基准,通过FlowFactor数据集发现不同因素下模型性能与现实准确性的关联,指出合成基准测试对现实数据准确性预测能力弱,扩大训练数据量未必能解决差距。
中文摘要 AI 辅助
将视觉模型进行现实世界部署以广泛造福社会是主要研究目标。然而在光流领域,由于难以获得地面真值,研究主要集中在合成数据和特定领域基准测试上。本文研究这种不匹配的严重程度,构建现实世界评估基准,用标准检查点评估一系列光流模型的现实世界通用性。基准包含TAP-Flow、Slow Flow及作者自建的FlowFactor数据集的8204帧对。FlowFactor含1000对高清帧,按大位移、重复纹理、遮挡和光照变化四个混杂因素组织。研究发现不同光照和大位移下的性能与现实世界准确性关联最强,大运动场景的改进可能影响小运动、静止场景的鲁棒性。实验表明在Sintel、KITTI和Spring上的进展对现实世界数据准确性预测能力弱,扩大训练数据量不一定能解决差距。
英文摘要
Real-world deployment of vision models to broadly benefit society is arguably a main research objective. In optical flow, however, the difficulty to obtain the ground truth has focused research mainly on synthetic data and domain-specific benchmarks. Here, we investigate the severity of this mismatch. We study how well modern optical flow estimation models generalise to real-world video and question if accuracy on synthetic benchmark proxies actually predicts accuracy on real-world optical flow. To address this, we build a real-world evaluation benchmark and evaluate the real-world generalisability of a broad set of recent optical flow models using standard checkpoints. Our benchmark contains 8,204 frame pairs across TAP-Flow, Slow Flow, and our own dataset FlowFactor. FlowFactor is a manually annotated real-world benchmark of 1,000 HD frame pairs organised into four confounding factors: large displacements, repetitive textures, occlusions, and lighting variation. Each setting mainly varies only one factor, enabling diagnostic, confounder-specific analysis. Using FlowFactor, we reveal that performance on varying lighting and large displacements correlates most strongly with real-world accuracy, and that improvements on large-motion regimes can trade off against robustness in small-motion, stationary scenes. Our experiments show that progress on Sintel, KITTI and Spring only weakly predicts accuracy on real-world data, highlighting the need for a broad real-world optical flow benchmark. Interestingly, scaling up the amount of training data does not necessarily resolve the gap, calling for new innovative research instead of simply scaling data and compute.
发表机构
- TU Delft(代尔夫特理工大学)
- Computer Vision Lab(计算机视觉实验室)
机构由 AI 辅助整理,请以论文原文为准。