AI 中文总结
研究视频模型能否从视觉提示工程中受益,通过自动修改任务图像提升性能,如视觉物理推理任务。发现视觉提示工程(VIPE)能提高视频推理性能,比传统方法更有效,为提升视频模型视觉推理性能提供简单高效途径。
AI 中文摘要
在基础模型时代,模型的性能取决于其提示。因此,提示工程已成为提高语言模型性能的关键技术。鉴于视频模型正成为视觉任务(如图像推理)的基础模型,本文探讨它们是否能从视觉提示工程中受益:自动修改任务图像以提升模型性能。例如,对于视觉物理推理任务,通过调用图像编辑模型可将抽象草图场景转化为逼真版本。研究发现,视觉提示工程(简称VIPE)能提升视频推理性能,且比传统文本提示工程或测试时缩放更有效。视觉提示工程可作为一种简单且计算高效的方法,从视频模型中引出更好的视觉推理性能。
英文摘要
In the age of foundation models, a model is only as good as its prompt. For this reason, prompt engineering has become an essential technique for improving language model performance. Since video models are currently becoming foundation models for visual tasks (e.g., visual reasoning), we here ask whether they similarly benefit from visual prompt engineering: automatically modifying the task image to improve model performance. For example, for a visual physics reasoning task ("Where does the ball land, after passing a set of obstacles?"), an abstract sketch-like scene can be turned into a photorealistic version with a simple call to an image editing model. We find that visual prompt engineering, or VIPE for short, improves video reasoning performance across tasks. In fact, for video models, visual prompt engineering can be even more effective than classic text-based prompt engineering or test-time scaling. Ultimately, just as text-based prompt engineering systematically improves language model performance, visual prompt engineering can serve as a simple, compute-efficient approach to elicit better visual reasoning performance from video models. Example videos on our project page at https://visual-prompt-engineering.github.io/.