arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

超越提示:将开发者的提问、行为与理解同编程智能体联系起来

Beyond the Prompt: Linking What Developers Ask, Do, and Understand with Coding Agents

Yunhan Qiao, Summit Haque, Christopher Hundhausen

arXiv 2609.33712首次发表:更新:

发表机构

Oregon State University, School of Electrical Engineering and Computer Science(俄勒冈州立大学电气与计算机工程学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出“说、做、理解”工作流程,通过提示、屏幕活动和事后解释分析开发者与编程智能体的交互,并在十名开发者使用GitHub Copilot的研究中识别出四种人物画像,揭示了任务表现与理解之间的分歧。

AI 中文摘要

编程智能体现在可以为开发者修改代码,开发者描述目标、提供上下文并响应智能体的工作。然而,提示、屏幕活动以及任务成功各自只能讲述这个故事的一部分。我们提出了“说、做、理解”(Say, Do, Understand),一个端到端的工作流程,用于分析开发者向智能体写了什么、在智能体工作时他们做了什么、以及之后他们能解释什么。该工作流程包含五个阶段(捕获、准备、分析、整合和解释)和三种工具:一个提示编码手册、一个用于对屏幕录制活动和事件进行编码的方案,以及分别用于解释过程和解决方案的评分细则。我们在一个观察性研究中应用了它,该研究涉及十名经验丰富的开发者,他们在不熟悉的代码库上使用GitHub Copilot。将任务表现与理解程度交叉产生了四种人物画像。两项衡量指标对八名开发者一致,但对两名开发者出现分歧:一人通过了大部分测试但无法解释解决方案,另一人通过了少量测试但解释得很好。在这个样本中,最常要求智能体检查其工作的人物画像在自己测试上花费的时间最少。这些模式是描述性的,不能推广到样本之外。我们向计算教育工作者、行业从业者和研究人员推荐该工作流程,以调整和评估软件工程中的人机交互。

英文摘要

Coding agents can now change code for developers, who describe goals, supply context, and respond to the agent's work. Yet prompts, screen activity, and task success each tell only part of this story. We present Say, Do, Understand, an end-to-end workflow for analyzing what developers write to an agent, what they do while it works, and what they can explain afterwards. The workflow has five stages (Capture, Prepare, Analyze, Integrate, and Interpret) and three instruments: a prompt codebook, a scheme for coding screen-recorded activities and events, and separate rubrics for explaining the process and the solution. We applied it in an observational study of ten experienced developers who used GitHub Copilot on an unfamiliar codebase. Crossing task performance with understanding produced four personas. The two measures agreed for eight developers but split for two: one passed most tests but could not explain the solution, and another passed few tests but explained it well. In this sample, the personas that most often asked the agent to check its work spent the least time testing on their own. These patterns are descriptive and do not generalize beyond the sample. We recommend the workflow to computing educators, industry practitioners, and researchers to adapt and evaluate human--AI communication in software engineering.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑