用于分层人类行为识别的组合式基准合成
Compositional Benchmark Synthesis for Hierarchical Human Action Recognition
查看机构详情
- University of Paris-Est Créteil (UPEC)(巴黎东克雷泰伊大学(UPEC))
- LISSI Laboratory(LISSI实验室)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
本文提出组合式基准合成框架,从单标签动作语料库合成四级分层人类行为识别基准,解决循环监督风险,经实验验证其结构属性并公开相关组件。
中文摘要 AI 辅助
识别从原子动作到长程意图的不同抽象层级的人类行为,需要带有语义层级标注的数据。大型语料库提供了孤立的、带有原子级标注的片段,但缺乏时间组合性;而记录复合活动的语料库则提供了浅层、领域狭窄且固定的层级结构。本文提出了一个基准生成与评估框架,该框架从平面单标签动作语料库中合成了一个包含动作、活动、低级意图(LLIs)和高级意图(HLIs)的四级分层意图基准,同时保留了动作层面真实的预提取特征。片段在主体一致性约束下通过转移模型组装而成,覆盖感知采样器将主体使用的基尼系数从0.566降至0.248。合成此类基准会产生记录数据集所避免的循环监督风险:若生成片段的规则也用于评估,模型可通过恢复生成器而非真正推理来取得成功。本文通过设计解决了有效性问题,将序列生成规则与评估所用的一阶逻辑规则分离。实例化后得到15002个片段。来自不同模型家族的四个参考基准用于表征难度,而非作为识别方法。所有基准均出现0.13至0.17的宏F1值的组合式保留差距,包括表现最佳的图感知模型也未缩小该差距,表明这是基准的结构属性而非模型人工产物。无逻辑基准在内在数据率之上仍违反保留的语义规则,破坏顺序的控制会在种子变异内改变宏F1值,用作生成器一致性检查。本体、转移模型和生成器已公开,以便可重新生成和扩展该基准。
英文摘要
Recognizing human behavior across levels of abstraction, from atomic actions to long-horizon intentions, requires data annotated along a semantic hierarchy. Large corpora provide isolated, atomically labeled clips without temporal composition, whereas recorded composite-activity corpora offer shallow, domain-narrow, fixedhierarchies. A benchmark-generation and evaluation frameworkis proposed that synthesizes a four-level hierarchical-intention benchmark, spanning actions, activities, low-level intentions (LLIs), and high-level intentions (HLIs), from a flat single-label action corpus while retaining real pre-extracted features at the action level. Episodes are assembled by a transition model under a subject-consistency constraint, and a coverage-aware sampler reduces the subject usage Gini from 0.566 to 0.248. Synthesizing such a benchmark raises a circular-supervision risk that recorded datasets avoid: if the rules generating the episodes also govern the evaluation, models can succeed by recovering the generator rather than through genuine reasoning. Validity is addressed by design, holding sequence-generation rules disjoint from the first-order-logic rules used at evaluation. The instantiation yields 15,002 episodes. Four reference baselines from different model families characterize difficulty, not as recognition methods. A compositional held-out gap of 0.13 to 0.17 macro-F1 appears across all baselines, including a graph-aware model that recognizes best yet does not close the gap, indicating a structural property of the benchmark rather than a model artifact. A logic-free baseline still violates the held-out semantic rules above their intrinsic data rate, and the order-destroying control changes macro-F1 within seed variation, serving as a generator-consistency check. Theontology, transition model, and generator are released so the benchmark can beregenerated and extended.