When Images Look Right and Retrieve Wrong: Coverage-Guided Cross-Scale Re-Indexing for Knowledge-Faithful Generative Perception
当生成图像看起来正确但检索错误:面向知识保真生成感知的覆盖度引导跨尺度重索引
机构 * National University of Singapore(新加坡国立大学) ; Wuhan University(武汉大学) ; ByteDance(字节跳动) ; Nankai University(南开大学) ; JD Research, JD.com, Inc.(京东研究院(京东有限公司)) ; Zhejiang University(浙江大学) ; King Abdullah University of Science and Technology(阿卜杜拉国王科技大学) ; The University of Hong Kong(香港大学)
AI总结 针对生成图像因尺度差异导致的语义崩溃问题,提出闭环多模态索引框架CERES,在多基准上实现SOTA,显著提升概念-查询检索性能。
Comments 20 pages, 7 figures, and 20 tables