arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于基于文本的人员异常搜索的动作对齐检索与成对多模态重排序

Action-Aligned Retrieval with Pairwise Multimodal Reranking for Text-Based Person Anomaly Search

Thanh-Khoi Nguyen, Thanh-Nhan Vo, Trong-Thuan Nguyen, Minh-Triet Tran

arXiv 2608.23503首次发表:更新:

发表机构

University of Science, VNU-HCM; Vietnam National University, Ho Chi Minh City(胡志明市国家大学科学大学; 胡志明市越南国家大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对基于文本的人员异常搜索中现有方法的局限,提出ActPair三阶段由粗到细框架,结合动作对齐检索与成对多模态重排序,在PAB测试集取得最优结果且可有效迁移至新数据集。

AI 中文摘要

基于文本的人员异常搜索需要依据细粒度、依赖上下文的行为而非单纯外观来区分个体。现有方法难以捕捉这些受上下文约束的动作,常依赖孤立的骨骼几何特征、在重构过程中丢弃原始查询细节,或使用绝对逐点评分进行多模态验证。为解决这些局限,我们提出ActPair,这是一个统一的三阶段由粗到细框架,结合动作对齐检索与成对多模态重排序以弥合姿态-语义鸿沟。首先,我们用动作对齐多任务目标微调视觉语言模型(VLM),该目标鼓励表征编码动作判别语义。其次,我们使用原始查询和大型语言模型(LLM)生成的基于上下文的改写进行并行后期融合检索,保留两种语义视图的互补细节。最后,我们提出一个高效的现成重排序模块,利用枢轴提升算法执行直接成对视觉比较,缓解残留的空间和组合歧义,且无穷举评估的高昂推理成本。大量实验表明,我们的框架在行人异常行为(PAB)公开测试集上取得了对比方法中的最佳结果,且能有效迁移到未见过的非异常特定数据集。

英文摘要

Text-based person anomaly search requires distinguishing individuals based on fine-grained, context-dependent behaviors rather than mere appearance. Existing methods struggle to capture these context-conditioned actions, frequently relying on isolated skeletal geometry, discarding raw query details during reformulation, or utilizing absolute pointwise scoring for multimodal verification. To address these limitations, we propose \textbf{ActPair}, a unified three-stage coarse-to-fine framework that combines action-aligned retrieval with pairwise multimodal reranking to bridge the pose-semantic gap. First, we fine-tune a vision-language model (VLM) with an action-aligned multi-task objective that encourages the representations to encode action-discriminative semantics. Second, we perform parallel late-fusion retrieval using the original query and a large language model (LLM)-generated context-grounded rewrite, retaining complementary details from both semantic views. Finally, we propose an efficient off-the-shelf reranking module that leverages a pivot-promote algorithm to perform direct pairwise visual comparisons, mitigating residual spatial and compositional ambiguities without the prohibitive inference costs of exhaustive evaluation. Extensive experiments demonstrate that our framework achieves the best results among the compared methods on the Pedestrian Anomaly Behavior (PAB) public test and transfers effectively to an unseen, non-anomaly-specific dataset.

CommentsAccepted to the AI City workshop @ ECCV 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑