arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.24127cs.CVcs.AI

Action-Slot:面向多智能体原子活动理解的结构化动作中心表示学习

Action-Slot: Structured Action-Centric Representation Learning for Multi-Agent Atomic Activity Understanding

Yu-Ho Chang, Chi-Hsi Kung, Yi-Hsuan Tsai, Yi-Ting Chen

中文总结 AI 辅助

本文提出Action-Slot框架,通过类别对齐动作槽、并行时空更新和背景正则化实现多智能体原子活动的结构化分解,并引入注意力差异伪掩码选择实现弱监督定位,在多个数据集上取得最优性能。

中文摘要 AI 辅助

原子活动理解旨在识别和定位结构化的交通行为,这些行为共同编码了运动模式及其在道路拓扑中的接地。与传统动作识别不同,原子活动是多智能体、多标签且拓扑感知的:多种活动同时发生,而许多智能体保持不活跃。我们引入了Action-Slot,一种结构化的动作中心表示学习框架。槽注意力广泛用于对象中心分解,但其置换不变设计和对象级归纳偏置与原子活动语义不一致。我们将槽学习重新表述为结构化活动分解,通过三种设计:(1)类别对齐的动作槽,将槽锚定到预定义的活动类别;(2)并行时空槽更新,用于整体视频级推理;(3)背景和负槽正则化,强制前景活动与不相关区域之间的竞争。这些共同建立了一种活动中心的归纳偏置,直接从原始视频中解开并发和异步活动。除了识别,学习到的表示编码了可迁移的时空接地信号。我们进一步提出了一种基于注意力差异的伪掩码选择框架,通过测量候选区域移除前后的注意力变化来抑制假阳性,从而实现无需密集标注的弱监督定位。为了支持系统评估,我们引入了TACO,一个平衡的合成数据集,具有完整的原子活动覆盖和像素级标注。在OATS、TACO和带标注的nuScenes上的实验显示了优越的识别性能、强大的模拟到真实迁移以及最先进的弱监督定位。

英文摘要

Atomic activity understanding aims to recognize and localize structured traffic behaviors that jointly encode motion patterns and their grounding in road topology. Unlike conventional action recognition, atomic activities are multi-agent, multi-label, and topology-aware: multiple activities co-occur while many agents remain inactive. We introduce Action-Slot, a structured action-centric representation learning framework. Slot attention is widely used for object-centric decomposition, but its permutation-invariant design and object-level inductive bias are misaligned with atomic activity semantics. We reformulate slot learning as structured activity decomposition through three designs: (1) category-aligned action slots that anchor slots to predefined activity categories, (2) parallel spatio-temporal slot updating for holistic video-level reasoning, and (3) background and negative-slot regularization that enforces competition between foreground activities and irrelevant regions. Together these establish an activity-centric inductive bias that disentangles concurrent and asynchronous activities directly from raw video. Beyond recognition, the learned representations encode transferable spatio-temporal grounding signals. We further propose an attention-difference-based pseudo mask selection framework that suppresses false positives by measuring attention changes before and after candidate region removal, enabling weakly supervised localization without dense annotations. To support systematic evaluation, we introduce TACO, a balanced synthetic dataset with full atomic activity coverage and pixel-level annotations. Experiments on OATS, TACO, and annotated nuScenes show superior recognition, strong sim-to-real transfer, and state-of-the-art weakly supervised localization.

补充信息

↑