发表机构
Artificial Intelligence Laboratory (Leibniz)(莱布尼茨人工智能实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对操作标注的自然语言难验证、刚性模板过度分割问题,提出SSC结构化表示,在BEHAVIOR-1K上验证其逻辑与内容补全功能,评估13个VL模型并报告标注异常。
AI 中文摘要
子任务标签将长时程操作演示分解为更短的语义片段,用于策略训练与评估。自然语言描述可读性强,但语言的变异性使其难以自动验证;而刚性模板格式(如BEHAVIOR-1K的skill_annotation)则存在语言过度分割问题,阻碍了可读性与标注一致性。我们提出结构化子任务链(Structured Subtask Chain,SSC),一种状态转换表示,可在上述两类极端方案间架起桥梁。一次演示由一系列结构化子任务模板(Structured Subtask Template,SST)条目构成,每个SST存储核心动作组件(主体、谓词、客体)、灵活条件(如空间或工具短语等状语修饰语)、与手臂动作分离的基础运动场,以及后状态场景图。基于该格式,SSC支持三种视觉语言辅助功能:将SST渲染为自然语言、依据四条状态转换规则校验组装后的链、通过查询解析级联补全未明确的字段。我们在BEHAVIOR-1K(含50个任务,每个任务3个演示片段,共2357个标注动作单元)上实例化该流水线,开展逻辑验证与内容补全,评估13个选定的先进视觉语言(VL)模型作为候选验证器,并报告标注异常情况。
英文摘要
Subtask labels decompose a long-horizon manipulation demonstration into shorter semantic segments for policy training and evaluation. Natural language descriptions are easy to read, but their linguistic variability makes automatic verification difficult. Rigid template formats, such as BEHAVIOR-1K's skill_annotation, are linguistically over-segmented, hindering both readability and annotation consistency. We propose the Structured Subtask Chain (SSC), a state-transition representation that bridges these extremes. A demonstration is a sequence of Structured Subtask Template (SST) entries. Each SST stores core action components (subject, predicate, object), flexible conditions (adverbial modifiers such as spatial or instrumental phrases), a base-motion field separate from arm actions, and an after-state scene graph. Built on this format, SSC supports three vision-language assisted functions: rendering SSTs as natural language, checking the assembled chain against four state-transition rules, and completing underspecified fields through a query resolution cascade. We instantiate the pipeline on BEHAVIOR-1K (50 tasks, 3 episodes per task, 2,357 annotated action cells) for logic verification and content completion, evaluating 13 selected state-of-the-art VL models as candidate verifiers and reporting labelling anomalies.
Comments8 main pages