发表机构
Vietnamese-German University; Ho Chi Minh City University of Technology; University of Information Technology; Ho Chi Minh city University of Science(越南-德国大学; 胡志明市技术大学; 信息技术大学; 胡志明市科学大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对Sim2Real基于文本的行人异常搜索的挑战,提出FaLCon框架,结合全局匹配与细粒度验证,在PAB基准上取得优异性能,代码将开源。
AI 中文摘要
基于文本的行人异常搜索需要使用主要在合成数据上训练的模型,从详细的自然语言描述中检索真实世界的行人图像。这种Sim2Real设置特别具有挑战性,因为视觉上相似的候选者可能仅在细微的动作、物体交互或外观属性上存在差异,而将多模态大语言模型应用于整个图像库的计算成本很高。我们提出了一种锚约束的粗到细检索框架,该框架将全局语义匹配与细粒度验证相结合。首先,每个查询由其原始标题、结构化连接和多个语义面表示。然后,通过鲁棒的逐查询分数校准和软声明感知融合来集成异构视觉-语言检索器。完整标题和连接标题用作锚点以保留候选召回,而外观、动作和物体面提供有界的校正证据。生成的候选池进一步通过判别式Qwen3重排序器以及两个互补的语义验证模块进行优化,这两个模块分别基于异常感知的完形填空和多智能体证据推理。最后,不确定性门控共识模块对模糊查询自适应地重新加权三个专家。在PAB基准上的实验表明,所提出的软声明感知检索达到了86.44%的mAP@10,显著优于单个检索骨干。完整框架进一步将性能提升至95.41%的mAP@10、94.44%的R@1和99.09%的R@5。这些结果表明,在将昂贵的语义推理限制在小候选池的同时保持强大的全局检索,对于细粒度Sim2Real行人异常搜索是有效的。我们的代码将在Github上提供。
英文摘要
Text-based person anomaly search requires retrieving real-world pedestrian images from detailed natural-language descriptions using models trained primarily on synthetic data. This Sim2Real setting is particularly challenging because visually similar candidates may differ only in subtle actions, object interactions, or appearance attributes, while applying multimodal large language models to the entire gallery is computationally expensive. We propose an anchor-constrained coarse-to-fine retrieval framework that combines global semantic matching with fine-grained verification. First, each query is represented by its original caption, a structured concatenation, and several semantic facets. Heterogeneous vision-language retrievers are then integrated through robust per-query score calibration and soft claim-aware fusion. Full and concatenated captions serve as anchors to preserve candidate recall, whereas appearance, action, and object facets provide bounded corrective evidence. The resulting candidate pool is further refined by a discriminative Qwen3 reranker and two complementary semantic verification modules based on anomaly-aware cloze completion and multi-agent evidence reasoning. Finally, an uncertainty-gated consensus module adaptively reweights the three experts on ambiguous queries. Experiments on the PAB benchmark show that the proposed soft claim-aware retrieval achieves 86.44% mAP@10, substantially outperforming individual retrieval backbones. The complete framework further improves performance to 95.41% mAP@10, 94.44% R@1, and 99.09% R@5. These results demonstrate that preserving strong global retrieval while restricting expensive semantic reasoning to a small candidate pool is effective for fine-grained Sim2Real person anomaly search. Our code will be available on Github.
Commentsaccepted to the ECCV 2026 AI City Challenge Workshop