DynaPix:视觉语言模型能否识别准确的未来?
DynaPix: Can Vision-Language Models Identify the Exact Future?
浏览论文内容
中文总结 AI 辅助
研究构建基准DynaPix,发现视觉语言模型能较好将预测与事件关联,但难以锚定时间,经模拟器真实记录训练可改善多数问题。
中文摘要 AI 辅助
在物理场景中行动需要知晓其真实的后续状态,而非看似合理的状态。当前评估通常接受文字或逼真的图像,因此预测状态从未与真实状态进行核对。我们引入DynaPix(动态像素)这一可核对预测结果的基准。给定一段在关键事件前停止的视频片段,以及关于后续时刻的问题,模型必须从相近候选或大型图库中挑选出真实的未来图像。这些场景来自物理模拟器,因此正确图像及其时间完全已知,而错误选项则被刻意设计得与真实图像相似。当可见事件标记目标时刻时,模型往往能成功;但当仅经过时间标记目标时刻时,模型表现接近随机。图库搜索难度更大,因为真实图像很少排在首位。人类能很好地处理经过时间标记的项目,因此难度在于模型而非问题。基于模拟器真实记录而非教师猜测的场景说明进行训练,可改善大部分问题,但无法解决更长时间间隔的情况。DynaPix由此揭示了时间锚定差距:模型将预测与事件关联的能力远优于与时间本身关联的能力。
英文摘要
Acting in a physical scene requires knowing its real later state, not a plausible one. Current evaluations often accept words or a realistic-looking image, so the predicted state is never checked against the true one. We introduce DynaPix (Dynamic Pixels), a benchmark that makes prediction checkable. Given a video clip that stops before a key event and a question about a later moment, a model must pick the true future image from close candidates or a large gallery. The scenes come from a physics simulator, so the correct image and its time are known exactly and the wrong options are deliberately similar. Models often succeed when a visible event marks the target moment, but are near chance when only elapsed time marks it. Gallery search is harder still, as the true image rarely ranks first. People handle the elapsed-time items well, so the difficulty lies with the models, not the questions. Training on scene accounts drawn from the simulator's true record, not a teacher's guess, repairs much of this but not the longer elapsed time case. DynaPix thus exposes a temporal-anchoring gap: models attach a prediction to an event far better than to time itself.
发表机构
- Centre for AI Research, VinUniversity(VinUniversity人工智能研究中心)
- National University of Singapore(新加坡国立大学)
机构由 AI 辅助整理,请以论文原文为准。