arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.04426cs.CVcs.CL

预测再检索:从视频前缀中进行跨实例未来状态检索

Predict, Then Retrieve: Cross-Instance Future-State Retrieval from Video Prefixes

Quynh Vo, Thong Nguyen, Vinh-Hien Do, Cong-Duy Nguyen, Anh-Tuan Luu

首次发表
浏览论文内容

中文总结 AI 辅助

该研究提出PSR任务,构建含多数据集的基准,提出LFTR模型,发现预测而非感知是核心挑战,LFTR缩小差距且成本低,相关资源已公开。

中文摘要 AI 辅助

我们提出了预测状态检索(Predictive State Retrieval, PSR)任务,在该任务中,模型观察一段短的视频前缀以及关于某物体未来状态的时间问题,然后从其他视频或图像中检索出描绘该状态的实例。与预测标签的动作预判、在视频内定位观察到事件的时刻检索、或合成像素的视频生成不同,PSR将预判与跨多个时间范围的跨实例检索相结合。我们从四个数据集构建了带有分级、经人工验证的真值、难度层级和理论上限的基准。我们还提出了LFTR,一种带有冻结编码器的轻量级检索器,它预测受问题和时间范围约束的未来隐变量,并在互补的语义和视觉空间中进行匹配。理论上限分解显示了一个明显瓶颈:一旦指定,真实未来状态极易检索,但我们评估的每个预测器,包括能访问前缀帧的大型多模态语言模型,仍远低于理论上限。因此,预测而非感知是核心可学习挑战。LFTR以显著更低的推理成本缩小了这一差距, ablation(消融实验)表明其增益源于跨空间融合和难负样本训练,而非隐变量展开。我们发布了该基准、代码和评估脚本。

英文摘要

We introduce Predictive State Retrieval (PSR), a task in which a model observes a short video prefix and a temporal question about an object's future state, then retrieves instances from other videos or images that depict that state. Unlike action anticipation, which predicts a label, moment retrieval, which localizes an observed event within a video, or video generation, which synthesizes pixels, PSR combines anticipation with cross-instance retrieval across multiple temporal horizons. We construct a benchmark from four datasets with graded, human-validated ground truth, difficulty tiers, and an oracle ceiling. We also propose LFTR, a lightweight retriever with frozen encoders that predicts a question- and horizon-conditioned future latent and matches it in complementary semantic and visual spaces. A ceiling decomposition reveals a clear bottleneck: the true future state is highly retrievable once specified, whereas every predictor we evaluate, including a large multimodal language model with access to the prefix frames, remains far below the oracle. Thus, forecasting rather than perception is the central learnable challenge. LFTR narrows this gap at substantially lower inference cost, and ablations attribute its gains to cross-space fusion and hard-negative training rather than latent rollout. We release the benchmark, code, and evaluation scripts.

发表机构

  • Centre for AI Research, VinUniversity(VinUniversity人工智能研究中心)
  • National University of Singapore(新加坡国立大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑