ParallelPilot:支持并行AI编码中的协调与监控
ParallelPilot: Supporting Coordination and Monitoring in Parallel AI Coding
- Microsoft Research AI Frontiers(微软研究院AI前沿部门)
- Columbia University(哥伦比亚大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
ParallelPilot通过规划界面、运行日志器和仪表板支持并行AI编码中的监督,提升票务吞吐量63%,并减少上下文切换,验证了显式监督支持的价值。
AI中文摘要:
随着编码助手变得越来越自主,开发者会并行运行多个会话,挑战从单纯的代码生成转变为协调和监控并发代理工作。通过一项形成性研究(N=14),我们确定了PILOT:用于规划、隔离、记录、观察和分流并行会话的五种监督实践。我们提出了ParallelPilot,一个设计探针,通过规划界面、运行日志器和环境仪表板,与现有编码工具一起实例化PILOT。在一项平衡的受试者内研究(N=16)中,使用ParallelPilot的参与者在短编码任务中票务吞吐量提高了63%,并在峰值时平均多监督一个并发代理,同时他们的跟踪努力和上下文切换有所减少。ParallelPilot还澄清了执行计划、任务依赖和干预线索,16名参与者中有14名更喜欢它而不是他们当前的设置。这些收益并未伴随感知控制或重定向代理感知成功的显著改善。我们的研究结果表明了显式监督支持的价值,并将PILOT定位为设计工具的支架,帮助人们在编码内外监督并发工作。我们建议未来的编码助手应将高层意识与低成本路径配对,回到开发者判断和引导代理工作所需的实现证据。
英文摘要:
As coding assistants become increasingly autonomous, developers run multiple sessions in parallel, shifting the challenge from code generation alone to coordinating and monitoring concurrent agent work. Through a formative study (N=14), we identified PILOT: five supervisory practices for Planning, Isolating, Logging, Observing, and Triaging parallel sessions. We present ParallelPilot, a design probe that instantiates PILOT through a planning interface, a run-logger, and an ambient dashboard alongside existing coding tools. In a counterbalanced within-subjects study (N=16), participants using ParallelPilot increased ticket throughput by 63% in short coding tasks and supervised an average of one more concurrent agent at peak, while their tracking effort and context switching dropped. ParallelPilot also clarified execution plans, task dependencies, and intervention cues, and 14 of 16 participants preferred it over their current setup. These gains were not accompanied by significant improvements in perceived control or perceived success in redirecting the agents. Our findings demonstrate the value of explicit supervision support and position PILOT as a scaffold for designing tools that help people supervise concurrent work within and beyond coding. We suggest that future coding assistants should pair high-level awareness with low-cost paths back to the implementation evidence developers need to judge and steer agent work.