arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.08216cs.AIcs.ETcs.LG

视觉:面向鲁棒且可解释智能体AI的数据中心锚定

Vision: Data-Centric Anchoring for Robust and Interpretable Agentic AI

Arun Vignesh Malarkkan, Xinyuan Wang, Yanjie Fu

首次发表
浏览论文内容

中文总结 AI 辅助

针对智能体AI在分布偏移下失效且不可解释的问题,提出数据中心锚定框架,通过策展、增强、约束、归因四阶段循环,将鲁棒性与可解释性设计进数据环境,实现自我纠正。

中文摘要 AI 辅助

基于大型语言模型的智能体AI系统存在两种持续性的失败模式,而这些是规模扩展无法解决的:它们在分布偏移下会失效,并且无法解释其做出的决策。我们认为,这些是管理智能体训练、评估和部署的数据生命周期中一个结构性缺陷的共同症状。观察性交互日志记录了智能体做了什么,而非它本可以做什么。它们编码了虚假相关性,缺乏受控变化,因此缺少区分因果信号与巧合或验证解释所需的反事实结构。没有模型中心的方法能恢复数据中从未包含的不变性。我们提出数据中心锚定:鲁棒性和可解释性应被设计进数据环境中,而非在训练后从模型中提取。我们的核心贡献是数据中心智能体循环,一个包含策展、增强、约束和归因的四阶段框架。这个顺序是结构性的,而非风格性的。策展先于增强,因为生成模型会放大其训练数据中的任何偏差。增强先于约束,因为没有跨环境的变异,不变性目标就是空洞的。归因闭环,将观察到的失败转化为下一轮迭代的有针对性的数据干预。每个阶段为下一阶段制造前提条件,这使得循环是自我纠正的,而非仅仅是顺序的。我们将该框架置于一个失败驱动的分类法中,将四种核心失败模式与数据生命周期联系起来:虚假特征依赖、分布偏移脆弱性、不确定性校准错误和解释不忠实。最后,我们讨论了该方法的局限性以及从当前状态到大规模实际部署之间存在的开放问题。

英文摘要

Agentic AI systems built on large language models fail in two persistent ways that scaling does not fix: they break under distribution shift, and they cannot explain the decisions they make. We argue these are co-symptoms of one structural deficiency in the data lifecycle that governs how agents are trained, evaluated, and deployed. Observational interaction logs record what an agent did, not what it would have done otherwise. They encode spurious correlations without controlled variation, so they lack the counterfactual structure needed to separate causal signal from coincidence or to validate an explanation. No model-centric method can recover invariances the data never contained. We present Data-Centric Anchoring: robustness and interpretability should be engineered into the data environment, not extracted from models after training. Our central contribution is the Data-Centric Agentic Loop, a four-stage framework of Curate, Augment, Constrain, and Attribute. The ordering is structural, not stylistic. Curation precedes augmentation because generative models amplify whatever bias they are trained on. Augmentation precedes constraint because invariance objectives are vacuous without variation across environments to be invariant to. Attribution closes the loop, converting observed failures into targeted data interventions for the next iteration. Each stage manufactures the preconditions of the next, which makes the loop self-correcting rather than merely sequential. We ground the framework in a failure-driven taxonomy that links four core failure modes to the data lifecycle: spurious feature reliance, distribution-shift fragility, uncertainty miscalibration, and explanation unfaithfulness. We close with the limits of this approach and the open problems that stand between it and practical deployment at scale.

发表机构

  • School of Computing and Augmented Intelligence, Arizona State University(亚利桑那州立大学计算与增强智能学院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑