arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于基于文本的行人异常检索的、带分歧感知重排序的异构视觉语言集成方法

Heterogeneous Vision-Language Ensemble with Disagreement-Aware Reranking for Text-Based Person Anomaly Retrieval

Huu-An Vu, Cam Tu Tran Thi, Thanh Toan Le Ngo, Hoang Vo, Do Trung Hieu, Hieu Dinh Trung Pham, Khang Minh Le, Huy Minh Nhat Nguyen

arXiv 2608.12843首次发表:更新:

发表机构

Hanoi University of Science and Technology; University of Information Technology, VNU-HCM; Vietnam National University, Ho Chi Minh City; VinUniversity; Vietnamese-German University(河内科技大学; 胡志明市国家大学信息技术大学; 胡志明市国家大学; VinUniversity; 越南德国大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对大规模基于文本的行人异常检索难题,提出带分歧感知重排序的异构视觉语言集成方法,在PAB基准上取得90.92% mAP等优异指标,验证了方案有效性。

AI 中文摘要

基于文本的行人异常检索旨在利用自然语言描述从大型图像库中检索表现出异常行为的行人。与传统的基于文本的行人检索相比,该任务需要对行人外观、行为、物体交互和场景上下文进行细粒度推理,这使得鲁棒的跨模态匹配极具挑战性。本文介绍GENAI4E团队对2026年AI City Challenge第4赛道的解决方案。我们的框架基于强大的检索骨干,通过分数对齐和迭代集成融合逐步整合异构视觉语言嵌入模型,随后对模糊查询进行分歧感知VLM重排序。在官方行人异常行为(Pedestrian Anomaly Behavior, PAB)基准上,我们的方法达到90.92%的mAP、85.13%的Recall@1、97.72%的Recall@5和98.68%的Recall@10,证明了结合互补视觉语言表示与选择性多模态推理用于大规模基于文本的行人异常检索的有效性。

英文摘要

Text-based person anomaly retrieval aims to retrieve pedestrians exhibiting anomalous behaviors from a large image gallery using natural language descriptions. Compared with conventional text-based person retrieval, this task requires fine-grained reasoning over pedestrian appearance, behaviors, object interactions, and scene context, making robust cross-modal matching significantly more challenging. This paper presents the GENAI4E team's solution to AI City Challenge 2026 Track 4. Our framework builds upon a strong retrieval backbone and progressively integrates heterogeneous vision-language embedding models through score alignment and iterative ensemble fusion, followed by disagreement-aware VLM reranking for ambiguous queries. On the official Pedestrian Anomaly Behavior (PAB) benchmark, our approach achieves 90.92% mAP, 85.13% Recall@1, 97.72% Recall@5, and 98.68% Recall@10, demonstrating the effectiveness of combining complementary vision-language representations with selective multimodal reasoning for large-scale text-based person anomaly retrieval.

CommentsAccepted at the ECCV 2026 Workshop

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑