arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

VIF-Bench:多参考图像生成中的视觉指令跟随评估

VIF-Bench: Evaluating Visual Instruction Following in Multi-Reference Image Generation

Yuta Oshima, Masakazu Yoshimura, Masahiro Suzuki, Yutaka Matsuo, Hiroki Furuta

arXiv 2609.37709首次发表:更新:

发表机构

The University of Tokyo(东京大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

VIF-Bench是一个包含1,241个任务的基准,用于评估多参考图像生成中模型在多个异构视觉指令下的遵循能力,发现遵循-伪影权衡、特定属性下的低遵循,并支持直接视觉约束优于文本描述。

AI 中文摘要

近期的多模态图像生成模型能够以多张图像和文本指令作为输入,实现不仅由文本引导,还由布局、箭头和姿态提示等视觉指令引导的基于参考的生成。然而,现有基准并未评估在多个异构视觉指令图像下组合多个参考的联合设置。为弥补这一空白,我们引入了VIF-Bench,一个包含1,241个任务的基准,旨在通过覆盖以下内容评估模型在该联合设置中的能力边界:(i) 在多个异构视觉指令(最多6个)下的多参考生成(最多7个参考),(ii) 参考图像可能与视觉指令产生竞争的情况(例如,强姿态主体与目标姿态),以及(iii) 在不同特异性水平下视觉指令与文本描述的受控比较。利用这些能力,我们发现了三个结论:(1) 模型面临遵循-伪影权衡:一旦模型达到更强的视觉指令遵循,更强的遵循往往与生成图像中更多的指令伪影同时出现;(2) 在参考图像携带受控属性的显著状态(例如,光照方向指令下的霓虹灯主体)的任务上,视觉指令遵循往往较低,这在光照和风方面最为一致;(3) 对于能够理解视觉指令的模型,直接提供视觉约束通常优于用文本描述它们;当使用文本时,适度的细节水平优于详尽描述。VIF-Bench作为开放基准发布,为可控多参考图像生成中的公平比较奠定基础。

英文摘要

Recent multimodal image generation models can take multiple images and textual instructions as input, enabling reference-based generation guided not only by text but also by visual instructions such as layouts, arrows, and pose cues. However, existing benchmarks do not evaluate the joint setting in which multiple references must be composed under multiple and heterogeneous visual-instruction images. To address this gap, we introduce VIF-Bench, a benchmark of 1,241 tasks designed to assess the edge of model capabilities in this joint setting by covering: (i) multi-reference generation (up to 7) under multiple heterogeneous visual instructions (up to 6), (ii) cases where reference images can potentially compete with visual instructions (e.g., a strongly posed subject vs. a target pose), and (iii) controlled comparison of visual instructions with text descriptions at different levels of specificity. Using these capabilities, we uncover three findings: (1) models face an adherence-artifact trade-off: once models reach stronger visual instruction adherence, stronger adherence tends to coincide with more instruction artifacts in generated images, (2) visual instruction adherence tends to be lower on tasks whose reference images carry a salient state of the controlled attribute (e.g., a neon-lit subject under a light-direction instruction), most consistently for light and wind, and (3) for models that can understand visual instructions, it is often better to provide visual constraints directly rather than describe them in text; when using text, a moderate level of detail works better than an exhaustive description. VIF-Bench is released as an open benchmark to establish a basis for fair comparison in controllable multi-reference image generation.

CommentsCode: https://github.com/shim0114/VIF-Bench , Benchmark: https://huggingface.co/datasets/shim0114/VIF-Bench

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑