arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.03554cs.CVcs.AI

WIDE:面向跨模态生成式检索的通配符动态扩展推理

WIDE: Wildcard Inference with Dynamic Expansion for Cross-Modal Generative Retrieval

Teng Guo, Xin Wang, Jiayou Xu, Keying Zhou, Jifeng Shen, Haoxin Ruan

AI总结:

本文针对跨模态生成式检索中因模态信息不对称导致的强制幻觉问题,提出WIDE方法,通过AET、AWD、BSR模块抑制幻觉,在M-BEIR基准上优于现有最优方法。

AI中文摘要:

生成式检索通过将表示学习与搜索统一为单个序列到序列的生成任务,已取得显著成功。然而,将该范式扩展至跨模态检索时,会因不同模态间固有的信息不对称性(如简洁文本查询与密集视觉候选间的差距)面临关键挑战。这种结构不匹配会导致自回归解码器在通过标准前缀树约束的束搜索生成标识符时出现强制幻觉,模型因无法猜测查询中缺失的细粒度细节而受到严重惩罚,使得无关候选者占据排名前列。为解决该问题,本文提出通配符动态扩展推理(Wildcard Inference with Dynamic Expansion,WIDE)。WIDE采用自适应熵阈值(Adaptive Entropy Thresholding,AET)离线校准各层的不确定性边界;在解码生成阶段,感知不对称性的通配符解码(Asymmetry-aware Wildcard Decoding,AWD)检测语义盲区并发射通配符以替代强制确定性标识符,动态扩展搜索空间且不产生对数概率惩罚;最后,盲区重排序(Blind-Spot Re-ranking,BSR)采用混合评分机制评估扩展后的候选池,该机制结合离散生成置信度与连续语义相似度。在M-BEIR基准上的大量实验表明,WIDE优于当前最优的生成式检索方法,可有效抑制强制幻觉,同时保持紧凑的索引结构。

英文摘要:

Generative retrieval has demonstrated significant success by unifying representation learning and search into a single sequence-to-sequence generation task. However, extending this paradigm to cross-modal retrieval reveals a critical challenge arising from the inherent information asymmetry across different modalities, such as the gap between concise text queries and dense visual candidates. This structural mismatch causes the autoregressive decoder to suffer from forced hallucination when generating identifiers via standard trie-constrained beam search, where the model is severely penalized for failing to guess fine-grained details absent from the query, allowing irrelevant candidates to hijack top rankings. To address this issue, we propose Wildcard Inference with Dynamic Expansion (WIDE). WIDE employs Adaptive Entropy Thresholding (AET) to calibrate layer-specific uncertainty boundaries offline. During the decoding generation phase, Asymmetry-aware Wildcard Decoding (AWD) detects semantic blind spots and emits wildcards instead of forced deterministic identifiers, dynamically expanding the search space without incurring log-probability penalties. Finally, Blind-Spot Re-ranking (BSR) evaluates the expanded candidate pool using a hybrid scoring mechanism that combines discrete generation confidence with continuous semantic similarity. Extensive experiments on the M-BEIR benchmark demonstrate that WIDE outperforms state-of-the-art generative retrieval methods, effectively suppressing forced hallucination while maintaining compact index structures.

补充信息

↑