arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

VISA:用于多模态指令遵循的智能体式自进化数据合成

VISA: Agentic Self-Evolving Data Synthesis for Multimodal Instruction Following

Min Zeng, Guanxin Tan, Libin Cen, Yafei Wen, Rui Hu, Liuyang Bian, Xiaolong Chen, Xiaoxin Chen

arXiv 2608.26013首次发表:更新:

发表机构

vivo AI Lab(vivo AI实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

VISA是一种智能体式自进化数据合成框架,通过自进化循环优化多模态指令合成,在MM-IFEval等基准上提升了多模态指令遵循性能。

AI 中文摘要

多模态指令遵循模型需要准确、多样、可验证且具有挑战性的训练数据。现有的合成流程通常遵循一次性生成-过滤范式,会丢弃失败样本、验证器结果以及目标模型错误带来的反馈。我们提出VISA(Visual Instruction Synthesis Agent,视觉指令合成智能体),这是一个将多模态指令合成重新表述为自进化循环的智能体框架。在每一轮中,VISA会分析图像以过滤不兼容的约束并发现新的可验证约束,从持久记忆中采样具有多样性和难度感知的约束集,生成候选指令,并使用可执行工具和结构化大语言模型判断器验证生成的样本。失败的样本会触发诊断引导的恢复,而被接受的样本会针对目标模型进行探测以评估难度。由此产生的验证器信号和目标模型失败轮廓会被写回记忆,使后续轮次能够自适应扩展约束空间、减少模板重复并聚焦于未解决的模型弱点。相同的验证器契约还为强化学习提供奖励信号,无需单独训练奖励模型。在MM-IFEval上的实验表明,VISA在强大的基线之上持续提升了多模态指令遵循性能,同时在七个公共基准上保留了通用多模态能力。

英文摘要

Multimodal instruction-following models require training data that is accurate, diverse, verifiable, and challenging. Existing synthesis pipelines typically follow a one-pass generate-and-filter paradigm, discarding feedback from failed samples, verifier outcomes, and target-model errors. We present VISA (Visual Instruction Synthesis Agent), an agentic framework that reformulates multimodal instruction synthesis as a self-evolving loop. At each round, VISA analyzes an image to filter incompatible constraints and discover new verifiable ones, samples diversity- and difficulty-aware constraint sets from persistent memory, generates candidate instructions, and verifies the resulting samples with executable tools and structured large language model judges. Failed samples trigger diagnostic-guided recovery, while accepted samples are probed against the target model to estimate difficulty. The resulting verifier signals and target-model failure profiles are written back to memory, allowing subsequent rounds to adaptively expand the constraint space, reduce template repetition, and focus on unresolved model weaknesses. The same verifier contracts further provide reward signals for reinforcement learning without a separately trained reward model. Experiments on MM-IFEval show that VISA consistently improves multimodal instruction following over strong baselines, while preserving general multimodal capability across seven public benchmarks.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑