arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

视觉世界的搜索:持久视觉记忆、分层索引与基于源的证据

Search over the Visual World: Persistent Visual Memory, Layered Indexes, and Source-Grounded Evidence

Sankalp Nagaonkar, Rohit Garg, Ankit Raj, Ashish Choithani, Ashutosh Trivedi

arXiv 2608.08075首次发表:更新:

发表机构

VideoDB(视频数据库)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对智能体持续观测视觉数据的场景,提出基于持久视觉记忆等的视频检索基础设施模型,经9800+查询对比,通用组件流水线在Recall@1/@3/@10上优于商业视频原生引擎

AI 中文摘要

大多数视频检索系统假设语料库是有界的,返回排名后的文件或时间戳。在摄像头、屏幕、流和档案上运行的智能体面临不同的系统问题:观测结果持续到达;模型以不同时间粒度对其进行解释;必须在不重放完整视觉记录的情况下选择上下文;且结果必须与可检查的源证据保持关联。我们认为,对这类语料库的搜索是一个基础设施问题,无法简化为对视频文件的排名。我们开发了一个针对视觉世界搜索的概念和形式模型,该模型基于分析器定义的场景、持久的理解人工制品、作为共享源时间上共存场景空间的视觉记忆,以及声明能力的索引,区分了记忆(保留的所有内容)、上下文(为任务选择的内容)和证据(为其提供基础的源区间)。VideoDB数据格式(VDB)在生产环境中实现了该模型,通过类型化搜索界面提供,涵盖计划检索、有状态调查、直接访问和基于源的合成。我们将这种与模型无关的基础设施(其中分割、采样、模型选择、嵌入和排名是系统决策,实时流是一等源)与作为固定API提供的视频原生基础模型进行对比。在针对商业视频原生引擎的语义检索对比中,使用四个公共数据集上的9800多个查询,通用组件流水线实现了更高的宏平均Recall@1/@3/@10(73.09/83.39/91.20,而商业引擎为65.75/77.13/89.10),而基线在Recall@50上更高(96.42,对比96.07)。如今,视觉世界的检索质量更多由系统设计而非视频特定预训练决定,视觉记忆基础设施可在将可播放的、基于源的证据作为一等要素的同时实现这一质量。

英文摘要

Most video-retrieval systems assume a bounded corpus and return ranked files or timestamps. Agents operating over cameras, screens, streams, and archives face a different systems problem: observations arrive continuously; models interpret them at different temporal granularities; context must be selected without replaying the complete visual record; and results must stay connected to inspectable source evidence. We argue that search over such a corpus is an infrastructure problem that cannot be reduced to ranking video files. We develop a conceptual and formal model of search over the visual world built on analyzer-defined scenes, persistent understanding artifacts, visual memory as coexisting scene spaces over shared source time, and capability-declared indexes, distinguishing memory (everything retained), context (what is selected for a task), and evidence (the source intervals that ground it). The VideoDB data format (VDB) realizes this model in production, exposed through a typed search surface spanning planned retrieval, stateful investigation, direct access, and grounded synthesis. We contrast this model-agnostic infrastructure, where segmentation, sampling, model choice, embeddings, and ranking are system decisions and live streams are first-class sources, with video-native foundation models offered as fixed APIs. In a semantic-retrieval comparison against a commercial video-native engine spanning 9,800+ queries over four public datasets, a pipeline of general-purpose components achieves higher macro-averaged Recall@1/@3/@10 (73.09/83.39/91.20 versus 65.75/77.13/89.10), while the baseline is higher at Recall@50 (96.42 versus 96.07). Retrieval quality over the visual world is today governed more by system design than by video-specific pretraining, and visual-memory infrastructure can deliver it while keeping playable, source-grounded evidence first-class.

Comments33 pages, 5 figures, 17 tables. Technical report. Benchmark configurations and reproduction instructions: https://github.com/video-db/search-over-the-visual-world

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑