AI 中文总结
该研究构建了名为TRACE的多层基准,向ALFRED轨迹注入受控漂移,经实验验证其可有效识别归因漂移,且重型注意力模型在该符号基准上无优势。
AI 中文摘要
现代网络物理系统与AI辅助系统将人类操作员、AI决策模块和自动化控制器耦合在单一控制回路中,其可信性取决于整个回路而非单个模型。然而,尚无标准基准能捕捉漂移与故障如何跨层传播的时间对齐多层轨迹,导致无法诊断协调失效的位置、原因及恢复方式。本文针对该缺口的一个方面——漂移(可源自任意堆栈层,传统单模态监控无法将其定位到具体层或确定发生时间)构建基准:向源自ALFRED(面向日常家务任务的接地指令基准)的轨迹中注入受控漂移,生成1918条漂移轨迹。每条轨迹为时间对齐的记录序列,涵盖5个执行层(状态、观测、决策、规则、控制),并标注漂移类型、受影响层、发生时间、责任主体及因果机制,且经独立评估者验证并报告了标注者间一致性。该数据集搭配了泄漏感知协议(消除近乎完美的发生时间泄漏),并开展了经典、循环及注意力型模型家族的基线研究。在该诚实协议下,所有模型家族均能以远高于随机和多数基线的水平识别并归因漂移(受影响层宏F1值近0.70,责任主体近0.85,因果机制近0.49),且在该符号基准上,重型注意力模型未比更简单模型展现优势。
英文摘要
Modern cyber-physical and AI-assisted systems couple human operators, AI decision modules, and automated controllers in a single control loop, so trustworthiness depends on the whole loop, not any one model. Yet no standard benchmark captures time-aligned, multi-layer traces of how drift and failures propagate across these layers, so we cannot diagnose where coordination breaks down, why, or how to recover. This paper targets one facet of that gap: drift, a deviation that can originate in any stack layer and that conventional single-modality monitoring cannot localize to a layer or pin to an onset time. We construct a benchmark by injecting controlled drift into traces derived from ALFRED, a grounded-instruction benchmark for everyday household tasks, yielding 1,918 drifted traces. Each trace is a time-aligned sequence of per-step records across five execution layers (state, observation, decision, rules, control), labeled with the drift type, affected layer, onset time, responsible actor, and causal mechanism, and validated by independent raters with inter-annotator agreement reported. We pair the dataset with a leak-aware protocol that removes a near-perfect onset leak, and a baseline study across classical, recurrent, and attention-based model families. Under this honest protocol, drift is identifiable and attributable well above random and majority baselines across every family (affected layer macro-F1 near 0.70, responsible actor near 0.85, causal mechanism near 0.49), and heavy attention offers no advantage over simpler models on this symbolic benchmark.
CommentsThis work was accepted for presentation and publication at the 25th IEEE International Conference on Machine Learning and Applications (ICMLA 2026)