arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.20810cs.MMcs.AIcs.CVcs.GR

当生成图像看起来正确但检索错误:面向知识保真生成感知的覆盖度引导跨尺度重索引

When Images Look Right and Retrieve Wrong: Coverage-Guided Cross-Scale Re-Indexing for Knowledge-Faithful Generative Perception

Guangyuan Dong, Chuang Liu, Haoyu Wang, Yangchen Zeng, Jiaqi Zhang, Li Jiuxing, Xiaoyang Yu, Pinlong Zhao, Yuchao Hou, Ziwei Li, Zheng Lin, Alexander Lim Han Yang, Yusen Wu

首次发表
浏览论文内容

中文总结 AI 辅助

针对生成图像因尺度差异导致的语义崩溃问题,提出闭环多模态索引框架CERES,在多基准上实现SOTA,显著提升概念-查询检索性能。

中文摘要 AI 辅助

多模态信息系统日益将生成的视觉内容反馈回支撑其生成的同一视觉-语言索引,因此生成结果必须能被其目标查询检索。当场景包含尺度差异极大的实体时,现有语言引导的生成器会基于单一全局池化文本嵌入进行条件生成,悄悄丢弃特定尺度的概念,即便像素保真度很高也会破坏概念-查询检索。我们将这种失效形式化为语义崩溃,提出CERES,这是一种闭环多模态索引框架,它构建三级语义金字塔,通过共现感知路由器挖掘隐式概念,向轻量型U-Net生成器执行尺度路由的交叉注意力,并通过用同一冻结的VLM对生成图像重索引来验证覆盖度。连续可微的软Jaccard覆盖度目标在明确的非退化条件下向参数规模为0.39M的生成器返回密集梯度,且覆盖度由仅在外部场景和对象标签上训练的独立DINOv2线性探针验证。在跨7个设置的4个 pansharpening基准上,CERES取得了新的SOTA,且在尺度变化最极端的场景中增益最大;相比最强基线,它还将概念-查询检索Recall@5提升了14.0个百分点,图像-文本平均倒数排名提升了0.19,表明该闭环保留了可查询内容,而非自指特征一致性。

英文摘要

Multimodal information systems increasingly route generated visual content back through the same vision-language index that informed its production, so the output must remain retrievable by the queries it was meant to serve. When the scene contains entities at vastly different scales, existing language-guided generators condition on a single, globally pooled text embedding and quietly drop scale-specific concepts, breaking concept-query retrieval even when pixel fidelity is high. We formalise this failure as semantic collapse and propose CERES, a closed-loop multimodal indexing framework that builds a three-level semantic pyramid, mines implicit concepts via a co-occurrence-aware router, performs scale-routed cross-attention into a lightweight U-Net generator, and verifies coverage by re-indexing the generated image with the same frozen VLM. A continuously differentiable soft-Jaccard coverage objective returns dense gradients to the 0.39 M-parameter generator under explicit non-degeneracy conditions, and coverage is verified by an independent DINOv2 linear probe trained only on external scene and object labels. On four pansharpening benchmarks across seven settings, CERES delivers the new state of the art with the largest gains where scale variation is most extreme (+4.64% relative Q2n and +9.7 mAP for DOTA detection). It also improves concept-query retrieval Recall@5 by +14.0 points and image-text mean reciprocal rank by 0.19 over the strongest baseline, showing that the closed loop preserves queryable content rather than self-referential feature consistency.

发表机构

  • National University of Singapore(新加坡国立大学)
  • Wuhan University(武汉大学)
  • ByteDance(字节跳动)
  • Nankai University(南开大学)
  • JD Research, JD.com, Inc.(京东研究院(京东有限公司))
  • Zhejiang University(浙江大学)
  • King Abdullah University of Science and Technology(阿卜杜拉国王科技大学)
  • The University of Hong Kong(香港大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑