arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于不确定性感知文本到视频检索的分布对齐桥

Distribution-Alignment Bridge for Uncertainty-Aware Text-to-Video Retrieval

Kyeongmo Chae, Jihoon Lee, Sangtae Ahn

arXiv 2607.20984首次发表:更新:

发表机构

School of Electronic and Electrical Engineering, Kyungpook National University(庆北国立大学电子与电气工程学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究文本到视频检索问题,提出分布对齐桥(DAB)框架,将其视为分布对齐任务,通过高斯分布建模考虑不确定性,用受扩散启发的桥和分布感知对比损失优化,在多基准测试中显著优于现有基线并提供校准排名。

AI 中文摘要

本文提出了分布对齐桥(DAB)框架,将文本到视频检索重新定义为分布对齐任务,而非传统确定性点匹配。通过将文本和视频嵌入建模为高斯分布,DAB明确考虑了模态特定的不确定性。利用确定性、受扩散启发的桥,通过截断细化过程迭代细化文本分布至目标视频分布。引入基于库尔贝克-莱布勒散度的分布感知对比损失优化跨模态相似性。在多个基准测试中,DAB显著优于现有基线,提供校准的不确定性感知排名。

英文摘要

This paper proposes the Distribution-Alignment Bridge (DAB), a framework that reconceptualizes text-to-video retrieval as a distribution alignment task rather than traditional deterministic point matching. By modeling both text and video embeddings as Gaussian distributions defined by mean and variance, DAB explicitly accounts for modality-specific uncertainty. We employ a deterministic, diffusion-inspired bridge to iteratively refine text distributions toward their target video distributions through a truncated refinement process. This approach unifies probabilistic embedding and distributional transformation into a cohesive, end-to-end trainable system. To optimize cross-modal similarity, we introduce a distribution-aware contrastive loss based on Kullback-Leibler divergence. Extensive evaluations on MSR-VTT, MSVD, and VATEX benchmarks confirm that DAB significantly outperforms existing probabilistic and diffusion-based baselines, while providing calibrated uncertainty-aware ranking through bridge-induced distributional margins.

CommentsECCV 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑