发表机构
Dalian University of Technology; University of International Business and Economics; Jinan University(大连理工大学; 对外经济贸易大学; 暨南大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出受人类视觉分层机制启发的EgoHieraLoc框架,含判别解析等三个模块及GSJC,在VQL-2D和VQL-3D基准上实现最优视觉查询定位性能。
AI 中文摘要
视觉查询定位(VQL)旨在从自我中心视频中检索并重新定位查询对象,但当对象边界模糊且全局上下文无法有效引导细粒度定位时,该任务仍具挑战性。人类视觉通过分层过程处理此类模糊性:快速筛选前景候选、在干扰项中选择性关注目标、通过全局上下文与局部细节间的反馈优化感知,且当单一视图不可靠时,会依据可信度整合多视角证据。受这些能力启发,我们提出EgoHieraLoc,这一适用于VQL-2D和VQL-3D的统一框架。判别解析模块首先利用分割先验提取感知前景的查询表示;查询感知模块随后通过带可变形建模的判别相关滤波实现鲁棒目标定位;区域适配模块将多尺度上下文反馈至局部区域,以恢复精确的对象边界。为将该感知分层扩展至3D定位,我们引入几何-语义联合置信度(GSJC),其将分割置信度与局部深度一致性、多视角反向投影一致性及三角测量基线质量相乘耦合,仅当视角在语义和几何上均可信时,才对3D估计做出贡献。大量实验表明,该方法在VQL-2D和VQL-3D基准上均取得了最优性能。
英文摘要
Visual query localization (VQL) aims to retrieve and re-localize a queried object in egocentric videos, yet remains challenging when object boundaries are ambiguous and global context cannot effectively guide fine-grained localization. Human vision handles such ambiguity through a hierarchical process: it rapidly screens foreground candidates, selectively attends to the target despite distractors, refines perception via feedback between global context and local detail, and, when a single view is unreliable, integrates evidence across viewpoints according to its credibility. Inspired by these competencies, we propose \textbf{EgoHieraLoc}, a unified framework for VQL-2D and VQL-3D. A Discriminative Parsing Module first extracts foreground-aware query representations using segmentation priors; a Query-Aware Module then performs robust target localization through discriminative correlation filtering with deformable modeling; and a Regional Adaptation Module feeds multi-scale context back into local regions to recover precise object boundaries. To extend this perceptual hierarchy to 3D localization, we introduce Geometric-Semantic Joint Confidence (GSJC), which multiplicatively couples segmentation confidence with local depth consistency, multi-view back-projection consistency, and triangulation-baseline quality, so that a viewpoint contributes to the 3D estimate only when it is credible both semantically and geometrically. Extensive experiments demonstrate state-of-the-art performance on both VQL-2D and -3D benchmarks.
Comments60 pages, under review