arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

指定和维护智能体工作流:GitHub 智能体工作流的实证研究

Specifying and Maintaining Agentic Workflows: An Empirical Study of GitHub Agentic Workflows

Jasem Khelifi, Issam Oukhay, Ali Ouni, Mohammed Sayagh, Mohamed Aymen Saied

arXiv 2609.27263首次发表:更新:

发表机构

École de technologie supérieure (ÉTS), University of Quebec; Université Laval(魁北克大学高等技术学院; 拉瓦尔大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过分析GitHub智能体工作流文件的结构、演变和指令内容,揭示了其长指令、持续更新及安全防护不足的特点,为开发者维护智能体工作流提供实证指导。

AI 中文摘要

智能体工作流将软件开发从针对单个任务提示人工智能智能体,转变为定义智能体自动执行的重复性工作。GitHub 智能体工作流(gh-aw)通过将自然语言指令与配置相结合的 Markdown 文件实现这种方法,这些文件可编译为可执行的 GitHub Actions 工作流。与传统主要规定脚本化操作的工作流不同,这些文件将需要解释的任务委托给人工智能智能体。它们还将智能体指令与执行触发器耦合,使这些指令成为重复性仓库活动的操作性规范。然而,开发者如何构建和维护这些规范,以及他们表达了哪些执行要求和保障措施,仍未被充分理解。在本文中,我们考察了 gh-aw Markdown 文件的结构、演变和指令内容,以指导从业者定义和维护智能体运行的工作。我们分析了来自 276 个仓库的 1,248 个文件、20,841 个提交-文件事件,以及从 294 个文件样本中得到的 288 个已解析的指令标签集。我们的结果表明,工作流指令超出了简短提示的范围,中位数为 556.5 个单词,62.1% 的文件包含代码块。在观察活动至少 120 天的文件中,78.2% 在第 4 个月仍收到更新,而规模归一化的变更量在第一个月后有所减少。任务、输出、约束和过程指令各自出现在超过 93% 的已标记工作流中,但只有 9.4% 明确涉及提示注入防御。LLM 分类对已解析的人工标签实现了 0.818 的 F1 分数和 0.715 的 Cohen's Kappa。这些发现表明,开发者应考虑跨仓库复制或引用的工作流的演变,并在适用时添加提示注入防御、资源预算和证据可信度检查。

英文摘要

Agentic workflows shift software development from prompting AI agents for individual tasks to defining recurring work that agents execute automatically. GitHub Agentic Workflows (gh-aw) enables this approach through Markdown files that combine natural-language instructions with configuration and compile into executable GitHub Actions workflows. Unlike conventional workflows that primarily prescribe scripted operations, these files delegate tasks requiring interpretation to AI agents. They also couple agent instructions with execution triggers, making those instructions operational specifications for repeated repository activities. However, how developers structure and maintain these specifications, and which execution requirements and safeguards they express, remains insufficiently understood. In this paper, we examine the structure, evolution, and instruction content of gh-aw Markdown files to inform how practitioners define and maintain agent-run work. We analyze 1,248 files from 276 repositories, 20,841 commit-file events, and 288 resolved instruction-label sets from a sample of 294 files. Our results show that workflow instructions extend beyond short prompts, with a median of 556.5 words and code blocks in 62.1% of files. Among files with at least 120 days of observed activity, 78.2% still receive updates in month 4, while size-normalized churn decreases after the first month. Tasks, outputs, constraints, and process instructions each appear in over 93% of labeled workflows, yet only 9.4% explicitly address prompt-injection defense. LLM classification achieves an F1 score of 0.818 and Cohen's Kappa of 0.715 against the resolved human labels. These findings suggest that developers should account for the evolution of workflows copied or referenced across repositories and consider adding prompt-injection defenses, resource budgets, and evidencecredibility checks where applicable.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑