发表机构
Universidad Politécnica de Madrid(马德里理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究基于AIDev数据集中的5,435个GitHub仓库,实证验证了智能体软件工程中八机制方法论约束框架的可观察性,发现机制采用以孤立形式为主,集成系统罕见。
AI 中文摘要
背景。智能体软件工程需要协调和管理智能体工作的机制。本研究框架提出了一种包含八种机制的方法论约束框架:上下文工程、持久共享知识、可执行规范、N版本思维与并行智能体、规范性规范、结构化咨询、基于证据的验收和渐进式自主。这些机制通过规则或上下文文件、规范和架构决策记录等工件实现,但该框架尚未在仓库规模上进行实证验证。目标。我们分析该约束框架在具有智能体活动的仓库中可观察到的程度和方式。RQ1通过工件流行度、广度、共现和时间演化来表征采用情况;RQ2检查规则文件——一种关键的可观察工件和智能体指令的持久来源——以评估其内容如何反映所提出的机制。方法。从包含116,211个GitHub仓库的AIDev数据集中,我们通过按可见性(以星标数作为流行度和/或声誉的指标)分层的设计分析了5,435个仓库。对于RQ1,我们检测并量化了七种可观察机制的工件和引入日期。对于RQ2,我们对来自150个仓库的规则文件进行了定性编码。结果。至少一种机制的人群流行率为21.7%,而在最可见的仓库中为65.3%。多机制配置很少见。在存在规则文件的地方,它们几乎总是指导智能体并陈述规范,而持久共享知识、可执行规范、结构化咨询和渐进式自主仅出现在少数仓库中。结论。该约束框架在经验上是可观察的,但主要通过孤立的机制而非框架所提出的集成系统实现。
英文摘要
Context. Agentic software engineering requires mechanisms to coordinate and govern agents' work. The framework motivating this study proposes a methodological harness with eight mechanisms: context engineering, persistent shared knowledge, executable specifications, N-version mindset and parallel agents, normative specifications, structured consultation, evidence-based acceptance, and graduated autonomy. These are realized through artifacts such as rule or context files, specifications, and architectural decision records, but the framework had not been empirically validated at repository scale. Objective. We analyze to what extent and how this harness is observable in repositories with agentic activity. RQ1 characterizes adoption through artifact prevalence, breadth, co-occurrence, and temporal evolution; RQ2 examines rule files - a key observable artifact and persistent source of agent instructions - to assess how their content reflects the proposed mechanisms. Method. From the AIDev dataset of 116,211 GitHub repositories, we analyzed 5,435 using a design stratified by visibility, measured through stars as an indicator of popularity and/or reputation. For RQ1, we detected and quantified artifacts and introduction dates for the seven observable mechanisms. For RQ2, we qualitatively coded rule files from 150 repositories. Results. Population prevalence of at least one mechanism is 21.7%, versus 65.3% among the most visible repositories. Multi-mechanism configurations are rare. Where rule files exist, they almost always guide the agent and state norms, while persistent shared knowledge, executable specifications, structured consultation, and graduated autonomy appear only in a minority. Conclusions. The harness is empirically observable, but mainly through isolated mechanisms rather than the integrated system proposed by the framework.