编织视觉叙事:超越原子视觉匹配的智能体图像束合成
Weaving Visual Narratives: Agentic Image Bundle Composition Beyond Atomic Visual Matching
浏览论文内容
中文总结 AI 辅助
针对传统图像检索原子匹配范式无法捕捉用户视觉搜索意图的问题,提出BundleWeaver智能体框架,构建IBCBench数据集,实现图像束合成并取得显著性能提升。
中文摘要 AI 辅助
图像检索传统上被表述为逐点匹配问题,其中每个候选图像被单独评分。然而,这种原子范式无法捕捉个人照片集中人类搜索意图的复杂性,用户通常寻求受结构关系约束的紧凑视觉故事,而非孤立的快照。为解决这一局限,我们提出图像束合成(Image Bundle Composition, IBC),这一新颖范式将目标从对单个图像排序转变为从海量非结构化照片池中动态合成内聚的图像束。由于目标束未预先定义,IBC面临严重的组合爆炸挑战,需要建模不可分解的联合相关性。为建立这一范式,我们构建IBCBench,首个IBC基准数据集,包含109467张图像和667个经验证的查询,通过半自动化验证流程构建。此外,我们提出BundleWeaver,一种智能体框架,将IBC重新表述为查询条件下的增量超边发现。通过利用大语言模型自适应搜索缺失的关系角色,并利用视觉语言模型进行整束验证,BundleWeaver有效探索组合空间。大量实验表明,尽管最先进的嵌入模型和静态分解重排序范式存在关系盲性,BundleWeaver实现了显著的性能提升,凸显了从原子评分转向动态关系合成的必要性。我们的数据集和代码可用。
英文摘要
Image retrieval has traditionally been formulated as a point-wise matching problem, where each candidate image is scored in isolation. However, this atomic paradigm fails to capture the complexity of human search intent within personal photo collections, where users often seek compact visual stories bound by structural relations rather than isolated snapshots. To address this limitation, we introduce **Image Bundle Composition (IBC)**, a novel paradigm that shifts the objective from ranking individual images to dynamically composing cohesive image bundles from a massive, unstructured photo pool. Since target bundles are not predefined, IBC presents a severe combinatorial explosion challenge and demands modeling non-decomposable joint relevance. To establish this paradigm, we construct **IBCBench**, the first IBC benchmark dataset containing 109,467 images and 667 verified queries, built via a semi-automated verification pipeline. Furthermore, we propose **BundleWeaver**, an agentic framework that reformulates IBC as query-conditioned incremental hyperedge discovery. By employing a Large Language Model to adaptively search for missing relational roles and utilizing a Vision-Language Model for whole-bundle verification, BundleWeaver effectively navigates the combinatorial space. Extensive experiments demonstrate that while state-of-the-art embedding models and static decompose-and-rerank paradigms suffer from relational blindness, BundleWeaver achieves substantial performance gains, highlighting the necessity of shifting from atomic scoring to dynamic relational composition. Our dataset and code are available.
发表机构
- Shanghai Jiao Tong University(上海交通大学)
- Shanghai Innovation Institute(上海创新研究院)
- OPPO(OPPO(广东欧珀移动通信有限公司))
机构由 AI 辅助整理,请以论文原文为准。