arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AgentForge:用于学习智能体软件工程的沉浸式角色扮演平台

AgentForge: An Immersive Role-Playing Platform for Learning Agentic Software Engineering

Zihan Fang, Yueke Zhang, Yu Huang

arXiv 2608.04148首次发表:更新:

发表机构

Vanderbilt University(范德堡大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出沉浸式角色扮演平台AgentForge,让新手在多智能体代码修复工作流中担任软件工程角色,经37名新手开发者测试,该平台可提升新手的软件工程技能与智能体协作理解。

AI 中文摘要

智能体人工智能正越来越多地被用于协调软件开发中的规划、实现、评审和测试工作,但其决策与交互过程的透明度往往有限。许多此类系统还假设用户能够有效引导AI的决策并验证其输出,这一假设对新手构成了特殊挑战——新手必须同时学习智能体AI的工作原理、如何与其有效协作,以及如何批判性地评估其输出。为应对这一挑战,我们提出了AgentForge,这是一个沉浸式学习系统,新手可在多智能体代码修复工作流中担任任务规划师、补丁作者、代码评审员或测试运行员四个软件工程角色之一。在每个实践环节中,新手履行所选角色职责,而AI智能体则承担其余三个角色的工作。AgentForge通过基于角色的支架式教学和元认知支持,明确各角色的特定职责,使智能体协调过程及中间产物可见,并鼓励新手监控和评估自身决策。在一项针对37名新手开发者的研究中,参与者在AI智能体的支持下实现了较高的任务完成率,但不同实践环节的交互需求存在显著差异:代码评审员实践环节需要更多的交互轮次、重定向操作和完成时间(校正后p值为0.004),且被认为是最具挑战性的。不过,参与者报告称其对软件修复和智能体协作的理解有显著提升(校正后p值小于0.001)。这些发现表明,AgentForge可帮助新手培养实用的软件工程技能,同时学习如何更批判性、更有效地与智能体AI协作。

英文摘要

Agentic AI is increasingly used to coordinate planning, implementation, review, and testing in software development, yet it often offers limited transparency into its decisions and interactions. Many such systems also assume that users can effectively guide the AI's decisions and validate its outputs. This assumption poses a particular challenge for novices, who must simultaneously learn how agentic AI works, how to collaborate with it effectively, and how to evaluate its outputs critically. To address this challenge, we present \textit{AgentForge}, an immersive learning system in which novices take on one of four software-engineering roles: Task Planner, Patch Author, Code Reviewer, or Test Runner, within a multi-agent code-repair workflow. In each practice session, the novices perform their chosen role while AI agents perform the remaining three. Through role-based scaffolding and metacognitive support, AgentForge clarifies role-specific responsibilities, makes agent coordination and intermediate artifacts visible, and encourages novices to monitor and evaluate their decisions. In a study with 37 novice developers, participants achieved high task-completion rates with AI-agent support. However, interaction demands differed significantly across practices: the Code Reviewer practice required more interaction turns, reroutes, and completion time ($p_{\mathrm{adj}} = .004$) and was perceived as the most challenging. Participants nevertheless reported significant gains in their understanding of software repair and agent collaboration ($p_{\mathrm{adj}} < .001$). These findings suggest that AgentForge can help novices develop practical software-engineering skills while learning to collaborate with agentic AI more critically and effectively.

Comments7 pages, 2 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑