深度之上的结构:用于多实体推理的单块时空变换器
Structure over Depth: A Single-Block Spatio-Temporal Transformer for Multi-Entity Reasoning
浏览论文内容
中文总结 AI 辅助
研究多实体时间数据建模问题,提出结构化时空变换器块,通过并行自注意力和双向交叉注意力及门控融合,在单个阶段明确建模三种交互,减少深层堆叠需求,在多任务中表现良好,支持结构优先设计原则。
中文摘要 AI 辅助
对多实体时间数据进行建模需要捕捉实体、时间及其交互之间的依赖关系。基于Transformer的方法表现良好,但通常依赖深层堆叠来隐式学习这些异构依赖关系,增加了计算成本。我们从结构角度重新审视这个问题,将多实体时间动态分解为三种交互类型:实体间的空间交互、跨时间的时间交互以及耦合两个域的交叉交互。我们提出了一个结构化的时空变换器块,在单个阶段中明确对所有三种交互进行建模。它使用并行的空间和时间自注意力,随后是双向交叉注意力,并通过可学习的门控融合组合输出。通过直接编码这些互补视图,该模型减少了对深层堆叠的需求。我们在基于视频的群体活动识别、基于骨架的人类交互分析和基于可穿戴传感器的活动识别上评估了该方法。尽管简单,但仅具有176万个参数的单个结构化Transformer块就能与更深的架构相匹配或超越它们。结果表明,先前模型中的深度部分补偿了隐式和纠缠的交互建模,但显式分解提供了一种更高效和透明的替代方案。更广泛地说,这项工作支持结构优先的设计原则:通过揭示交互结构而不是依赖深度,可以实现富有表现力的多实体时间推理。
英文摘要
Modeling multi-entity temporal data requires capturing dependencies across entities, time, and their interactions. Transformer-based approaches perform well but often rely on deep stacks of layers to learn these heterogeneous dependencies implicitly, increasing computational cost. We revisit this problem from a structural perspective and decompose multi-entity temporal dynamics into three interaction types: spatial interactions among entities, temporal interactions across time, and cross interactions coupling the two domains. We propose a structured spatio-temporal transformer block that explicitly models all three within a single stage. It uses parallel spatial and temporal self-attention, followed by bidirectional cross-attention, and combines the outputs through learnable gated fusion. By directly encoding these complementary views, the model reduces the need for deep stacking. We evaluate the approach on video-based group activity recognition, skeleton-based human interaction analysis, and wearable sensor-based activity recognition. Despite its simplicity, the single structured Transformer block matches or outperforms deeper architectures with only 1.76M parameters. The results suggest that depth in prior models partly compensates for implicit and entangled interaction modeling, whereas explicit factorization offers a more efficient and transparent alternative. More broadly, this work supports a structure-first design principle: expressive multi-entity temporal reasoning can emerge by exposing interaction structure rather than relying on depth.