发表机构
Hong Kong University of Science and Technology; University of Waterloo; Qwen Applications; Imperial College London(香港科技大学; 滑铁卢大学; 通义千问应用; 帝国理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对视觉生成器的世界知识瓶颈,构建专用数据集与基准,提出教-搜协同训练框架定位动态知识边界,实现知识驱动的可迭代视觉生成性能提升。
AI 中文摘要
视觉生成器擅长渲染,但会自信地生成自身并不了解的内容。用户请求是无界、动态演化且高度长尾的:涵盖全新角色、流行实体、训练截止日期之后发生的事件等各类内容。这种世界知识瓶颈是结构性的:生成器基于固定语料训练,而视觉世界是开放的。我们构建了包含20839条提示的SearchGen-20K数据集与SearchGen-Bench基准,覆盖12类失效场景与22个领域,配套预执行的多模态SearchGen-Corpus-1M语料库,为离线可复现研究提供支撑。在SearchGen-Bench上,当前顶尖的开源生成器得分仅为21至28分(满分100),这一40分的性能滑坡在现有基准中完全无法被观测到。自然的解决方案是引入搜索工具,实现智能体视觉生成。但我们发现,朴素搜索方案效果不佳:它会无差别检索内容,向生成器原本可以正确处理的提示中注入噪声。我们将根本原因定位为生成器特有的动态演化知识边界:即生成器可通过训练内化的内容与必须留存于外部上下文的内容之间的分界。尽管该边界难以预先明确指定,我们证明其可通过先教后搜的协同训练框架被探测定位。即使是该协同训练方案的最简版本也能带来性能的单调提升,为可满足世界知识锚定请求的视觉生成递归自优化能力奠定基础。我们公开全部数据集、协同训练语料与搜索语料,作为可复现的工具增强型、世界知识锚定视觉生成研究支撑框架。
英文摘要
Visual generators excel at rendering, but they confidently fabricate what they do not know. User requests are unbounded, evolving, and deeply long-tailed: new characters, trending entities, post-cutoff events, and more. This world-knowledge bottleneck is structural: generators are trained on fixed corpora, but the visual world is open-ended. We construct SearchGen-20K and SearchGen-Bench, with 20,839 prompts spanning twelve failure categories and twenty-two domains, paired with a pre-executed multimodal SearchGen-Corpus-1M to support offline, reproducible research. On SearchGen-Bench, frontier open generators score only 21 to 28 out of 100, a 40-point collapse invisible to existing benchmarks. The natural remedy is to employ search tools, enabling agentic visual generation. However, we find that naive search fails: it retrieves indiscriminately, injecting noise into prompts the generator already handles. We trace the root cause to a generator-specific, evolving knowledge boundary: the divide between what a generator can internalize through training and what must remain in external context. Although this boundary is hard to specify in advance, we show that it is discoverable through a teach-then-search co-training framework. Even a minimal version of this co-training recipe produces monotonic improvement, laying the foundation for recursive self-improvement in visual generation that can meet world-knowledge-grounded requests. We release the full dataset, co-training corpus, and search corpus as a replayable harness for tool-augmented, world-knowledge-grounded visual generation.