结构化时空证据图用于视频中的开放词汇对象检索
Structured Spatio-Temporal Evidence Graphs for Open-Vocabulary Object Retrieval in Videos
浏览论文内容
中文总结 AI 辅助
提出STEG-OVR结构化时空证据图,将视频对象检索从孤立框匹配升级为轨迹与关系证据匹配,显著提升停止状态和场景占用类查询的AP。
中文摘要 AI 辅助
视频中的开放词汇对象检索需要在有界查询时间成本下回答自由形式的对象查询。现有的基于索引的系统通常存储独立的帧级区域,并通过视觉-语言相似性进行检索,这种方法对外观查询有效,但与谓词不匹配,这些谓词的证据是时间性或关系性的,例如停止状态、场景区域占用、持续性和对象交互。我们将这一差距识别为证据单元不匹配:查询是在轨迹段或对象元组上表达的,而索引存储的是孤立的边界框。为解决这一问题,我们提出了STEG-OVR,一种用于开放词汇对象检索的结构化时空证据图。STEG-OVR将持久对象表示为轨迹段节点,将时间上兼容的对象对表示为关系边,存储外观、运动、场景占用、相对几何、速度兼容性和符号关系证据。查询被分解为实体、状态、场景、时间和关系槽位,这些槽位在软分数融合和固定预算一致性验证之前仅激活相应的检索通道。在三个以对象为中心的设置上的诊断实验显示,Beach上的AP从0.0882提高到0.2073,Shibuya上的AP从0.0834提高到0.1505,LOVO风格比较中的AP从0.701提高到0.743。增益在停止状态和场景区域占用查询上最强,而持续关系仍对跟踪连续性和谓词校准敏感。
英文摘要
Open-vocabulary object retrieval in videos requires answering free-form object queries under bounded query-time cost. Existing index-based systems typically store independent frame-level regions and retrieve them with vision-language similarity, which is effective for appearance queries but mismatched with predicates whose evidence is temporal or relational, such as stopped state, scene-region occupancy, persistence, and object interactions. We identify this gap as an evidence-unit mismatch: the query is expressed over tracklets or object tuples, while the index stores isolated boxes. To address it, we propose STEG-OVR, a structured spatio-temporal evidence graph for open-vocabulary object retrieval. STEG-OVR represents persistent objects as tracklet nodes and temporally compatible object pairs as relation edges, storing appearance, motion, scene occupancy, relative geometry, velocity compatibility, and symbolic relation evidence. A query is decomposed into entity, state, scene, temporal, and relation slots, which activate only the corresponding retrieval channels before soft score fusion and fixed-budget consistency verification. Diagnostic experiments on three object-centric settings show AP improvements from 0.0882 to 0.2073 on Beach, from 0.0834 to 0.1505 on Shibuya, and from 0.701 to 0.743 in a LOVO-style comparison. The gains are strongest for stopped-state and scene-region-occupancy queries, while sustained relations remain sensitive to tracking continuity and predicate calibration.
发表机构
- School of Computer Science and Engineering, Northeastern University(东北大学计算机科学与工程学院)
机构由 AI 辅助整理,请以论文原文为准。