arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

LightAIR:用于基于文本的行人异常搜索的轻量型动作逆转换与黎曼校正

LightAIR: Lightweight Action Inversion and Riemannian Rectification for Text-based Person Anomaly Search

Yulun Zhang, Zixu Li, Zhiwei Chen, Zhiheng Fu, Wenbo Wang, Zihang Qiu, Zhilin Wang, Ruxin Wang, Yupeng Hu

arXiv 2608.09152首次发表:更新:

发表机构

Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences; Shandong University(中国科学院深圳先进技术研究院; 山东大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对现有基于文本的行人异常搜索方法的缺陷,提出LightAIR网络,通过动作逆转换、正交零空间投影与黎曼梯度校正实现外观与动作特征解耦,在TPAS和TIPR数据集上性能优于现有SOTA方法

AI 中文摘要

传统的基于文本的行人搜索(TPS)通常局限于匹配静态外观属性,严重忽略了动态动作信息。基于文本的行人异常搜索(TPAS)任务填补了这一空白,要求模型在匹配行人宏观外观的同时定位微观层面的特定异常行为。然而,当前的TPAS方法存在根本性局限:外部显式姿态估计器在无约束监控场景中表现脆弱,而隐式学习在像素级纠缠下会遭遇视觉解耦失败,导致主导的外观信息极易掩盖并污染细微的动作特征。此外,在传统欧氏空间中对“相同外观、不同动作”的难负样本进行对比优化会引发严重的捷径学习。为解决这些问题,我们提出了轻量型动作逆转换与黎曼校正网络(LightAIR)。首先,它通过轻量型动作逆转换算子引入文本语义先验作为锚点,以提取纯动作特征,从而克服视觉固有耦合。随后,它采用正交零空间投影将外观特征约束在动作特征的正交补空间内,保证严格的前向解耦。最后,我们设计了梯度校正模块,该模块计算黎曼梯度以约束反向传播轨迹,迫使梯度流严格沿保留解耦特性的切空间更新,从而切断有害捷径。在广泛使用的TPAS和TIPR数据集上进行的大量实验表明,LightAIR显著优于现有的最先进方法。代码可在this https URL获取

英文摘要

Traditional Text-based Person Search (TPS) is typically limited to matching static appearance attributes, severely neglecting dynamic action information. The Text-based Person Anomaly Search (TPAS) task bridges this gap, requiring models to locate micro-level specific abnormal behaviors while matching macro-level appearance of pedestrians. However, current TPAS methods face fundamental limitations: external explicit pose estimators are fragile in unconstrained surveillance scenarios, and implicit learning encounters visual decoupling failure under pixel-level entanglement, causing dominant appearance information to easily swallow and contaminate subtle action features. Furthermore, performing contrastive optimization on hard negative samples (``same appearance, different actions'') in conventional Euclidean spaces induces severe shortcut learning. To address these, we propose the Lightweight Action Inversion and Riemannian rectification network (LightAIR). First, it introduces textual semantic priors as anchors via a lightweight action inversion operator to extract pure action features, thereby overcoming visual-inherent coupling. Subsequently, it employs orthogonal null-space projection to constrain appearance features within the orthogonal complement space of action features, guaranteeing strict forward decoupling. Finally, we designed a gradient rectification module that computes the Riemannian gradient to constrain the backpropagation trajectory, forcing the gradient flow to update strictly along the tangent space that preserves decoupling properties, thereby cutting off harmful shortcuts. Extensive experiments on the widely used TPAS and TIPR datasets demonstrate that LightAIR significantly outperforms existing state-of-the-art methods. Codes are available at https://github.com/rainy-london/LightAIR

CommentsAccepted by ACM MM 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑