arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ExpertVerse:知识密集型视觉合成中专家级推理的通用基准

ExpertVerse: A General-Purpose Benchmark for Expert-Level Reasoning in Knowledge-Intensive Visual Synthesis

Yuan Wang, Yongchao Du, Mengting Chen, Jinsong Lan, Jiaxuan Luo, Xuetao Feng, Xiaoyong Zhu

arXiv 2607.19341首次发表:更新:

AI 中文总结

研究针对多模态生成模型在知识密集型生成的不足,开发ExpertVerse基准,涵盖多种认知能力和专家学科,生成相关数据集,训练KnowThinker并提出BPPO优化方法,揭示模型推理缺陷,强调知识密集型基准对下一代视觉生成的重要性。

AI 中文摘要

多模态生成模型的进展使基于指令的图像生成从语义操作转向知识驱动的视觉推理,但这些方法在知识密集型生成方面存在不足。我们开发了ExpertVerse,一个以能力为中心的基准,通过知识密集型视角评估生成模型。它按9种认知能力和8个专家学科的正交分类法对推理生成进行分层,产生58个子学科。我们策划了1611个专家注释实例,还开发了自动化工作流程生成ExpertVerse-100K数据集。基于此训练了KnowThinker,并针对多奖励优化中的跨模态信用失调和多目标梯度冲突提出了BPPO。开源和专有模型的大量结果揭示了关键推理缺陷,凸显了知识密集型基准对下一代视觉生成的必要性。

英文摘要

Recent advances in multimodal generative models have enabled instruction-based image generation to move beyond semantic manipulation to knowledge-driven visual reasoning. However, these methods focus on explicit commonsense reasoning, shallow causal understanding, and direct knowledge recall, failing at knowledge-intensive generation. We develop \textbf{ExpertVerse}, a capability-centric benchmark to evaluate generative models via knowledge-intensive lens. ExpertVerse stratifies reasoning generation across an orthogonal taxonomy of \textit{9 cognitive capabilities} and \textit{8 expert disciplines}, yielding \textit{58 sub-disciplines}. We curate 1,611 expert-annotated instances covering single-image editing, multi-image composition, and text-to-image generation. We further develop an automated workflow to produce \textbf{ExpertVerse-100K}, a large-scale dataset with reasoning traces and knowledge-anchored rationale annotations. Based on this, we train \textbf{KnowThinker} with RL fine-tuning, a VLM reasoning engine with world knowledge that jointly generates thinking processes and refined instructions. Towards the cross-modal credit misalignment and multi-objective gradient conflicts in multi-reward optimization, we propose a tailored Bootstrapped Pareto Policy Optimization (BPPO), which synergizes Bootstrapping Reward Rectification (BRR) and Conflict-Aware Pareto Advantage Fusion (CPAF). Extensive results of both open-source and proprietary models exposes critical reasoning deficits, highlighting imperative for knowledge-intensive benchmarks towards next-generation visual generation.

Comments14 pages, 6 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑