arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.29604cs.IRcs.CV

RePair:将检索失败转化为反事实困难对

RePair: Turning Retrieval Failures into Counterfactual Hard Pairs

  • The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
  • Tencent Yuanbao(腾讯元宝)
  • The University of Hong Kong(香港大学)
  • ARC Lab, Tencent(腾讯ARC实验室)

机构由 AI 辅助整理,请以论文原文为准。

Siyi Liu, Xiaorong Zhu, Enjun Du, Xinyu Zuo, Lisheng Duan, Haijin Liang, Jin Ma, Junfu Pu, Yongqi Zhang

AI总结:

RePair基于CLIP双编码器的视觉-语言检索,通过挖掘假阳性并结合LLM引导的反事实编辑生成困难对,提升数据效率,在Flickr30K和COCO30K上性能优于基线。

AI中文摘要:

采用CLIP风格双编码器的视觉-语言检索已取得优异的跨模态性能,但实际准确率往往依赖于局部语义区分——排名靠前的近似匹配与真实匹配仅存在一处关键细节差异。困难样本挖掘可筛选出易混淆候选,但无法构造修正后的对应样本;合成增强可生成新样本,但若不基于模型实际失败情况,会针对无关的困难维度。我们发现,排名靠前的假阳性是一种反事实支架:它共享查询的大部分语义,仅在导致失败的局部残差上存在差异。对该残差进行最小限度修正,可得到与真实匹配同模态的困难正样本;修正后的版本与未编辑版本构成跨越决策边界的困难负样本,能提供互补的拉-推监督信号。我们提出RePair,其遵循有效性、最小性和局部性三条原则,双向挖掘假阳性,应用大语言模型(LLM)引导的反事实编辑,并采用局部困难对对比损失进行训练。在Flickr30K和COCO30K数据集上,RePair仅使用10.7万个合成样本,就优于受控增强基线,比同类方法少26%至75%的合成样本,证实基于失败的修复比不考虑错误的增强更具数据效率。

英文摘要:

Vision-language retrieval with CLIP-style dual encoders achieves strong cross-modal performance, yet practical accuracy often hinges on localized semantic distinctions where top-ranked near misses differ from the true match by a single critical detail. Hard-sample mining can select confusable candidates but cannot construct corrected counterparts; synthetic augmentation can generate novel samples but, without conditioning on actual model failures, targets irrelevant dimensions of hardness. We observe that a top-ranked false positive is a counterfactual scaffold---sharing most of the query's semantics while differing in a localized failure-causing residual. Minimally correcting this residual yields a hard positive of the ground truth in the same modality; the corrected and unedited versions form a hard negative pair that straddles the decision boundary, producing complementary pull--push supervision. We introduce RePair, guided by three principles---Validity, Minimality, and Locality---which mines false positives bidirectionally, applies LLM-guided counterfactual editing, and trains with a local hard-pair contrastive objective. On Flickr30K and COCO30K, RePair outperforms controlled augmentation baselines with only 107K synthetic samples---26\%--75\% fewer than comparable methods---confirming failure-conditioned repair is more data-efficient than error-agnostic augmentation.

补充信息

↑