arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.01955cs.ETcs.AIcs.DB

面向数据与AI流水线的智能体自修复:一种采用开源软件的低成本、厂商无关架构

Agentic Self-Healing for Data and AI Pipelines: An Affordable Vendor-Agnostic Architecture using Open-Source Software

Solomon Eshun, Dennis Murage, Sharleen Muoki, Chih-Chun Chen, Stephen Adjignon, Matteo Staar, Oliver Angélil

首次发表
浏览论文内容

中文总结 AI 辅助

本文对比现有AI辅助流水线运维方案的不足,提出采用开源工具的低成本厂商无关智能体自修复流水线架构,可帮助团队减少人工处理流水线问题的工作量。

中文摘要 AI 辅助

现代组织依赖数据、机器学习和软件交付流水线来传输数据、训练模型、部署应用、更新仪表盘并支撑业务关键决策。然而这些流水线常因数据质量问题、模式变更、上游数据源变更、基础设施问题、编排故障及模型工作流问题而失效。现有的零运维(ZeroOps)、可观测性及AI运维平台可帮助团队检测事件、调查根本原因,部分还能推荐或执行修复,但许多此类方案成本高昂、厂商特定,小型团队难以在不同工具和环境中适配。本文首先对比了现有AI辅助流水线监控、根本原因分析及自动修复的现成方案,包括其优势、局限性和实际权衡。基于此对比,我们发现主要差距在于架构而非技术:自修复流水线所需的要素已存在,但分散在厂商特定平台、可观测性工具、事件系统及开源组件中。因此,我们提出一种采用开源和低成本工具的、低成本、厂商无关的智能体自修复流水线参考架构。该架构结合监控、流水线元数据、事件历史、确定性策略检查、AI辅助诊断、审批工作流及受控修复行动,帮助团队以更少的人工工作量检测、诊断、修复、验证并从流水线问题中学习,目标是提供可在数据工程、机器学习运维及软件交付环境中适配的实用参考架构。

英文摘要

Modern organizations rely on data, machine learning, and software delivery pipelines to move data, train models, deploy applications, refresh dashboards, and support business-critical decisions. However, these pipelines often fail because of data quality issues, schema changes, upstream source changes, infrastructure problems, orchestration failures, and model workflow issues. Existing ZeroOps, observability, and AI operations platforms can help teams detect incidents, investigate root causes, and in some cases recommend or execute fixes. However, many of these solutions are expensive, vendor-specific, or difficult for smaller teams to adapt across different tools and environments. This paper first compares existing off-the-shelf solutions for AI-assisted pipeline monitoring, root-cause analysis, and automated remediation, including their strengths, limitations, and practical trade-offs. Based on this comparison, we find that the main gap is architectural rather than technological: the required ingredients for self-healing pipelines already exist, but they are fragmented across vendor-specific platforms, observability tools, incident systems, and open-source components. We therefore propose an affordable, vendor-agnostic reference architecture for agentic self-healing pipelines using open-source and low-cost tools. The proposed architecture combines monitoring, pipeline metadata, incident history, deterministic policy checks, AI-assisted diagnosis, approval workflows, and controlled remediation actions to help teams detect, diagnose, repair, verify, and learn from pipeline issues with less manual effort. The goal is to provide a practical reference architecture that can be adapted across data engineering, machine learning operations, and software delivery environments.

发表机构

  • Karlsruhe University of Applied Sciences(卡尔斯鲁厄应用技术大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑