WatchPoint:面向真实世界智能体化Web开发的可执行用户反馈
WatchPoint: Executable User Feedback for Real-World Agentic Web Development
浏览论文内容
中文总结 AI 辅助
WatchPoint通过模拟开发者行为生成并执行诊断脚本,为编码智能体提供可执行反馈,在Web-Bench上恢复57.6%的任务,与人类测试者54.5%的恢复率相当。
中文摘要 AI 辅助
当专业Web开发者的代码未通过测试时,他们不会仅仅重新阅读堆栈跟踪信息。他们会在浏览器中打开应用程序,点击按钮,检查计算样式,并运行诊断命令以理解问题所在。现有的编码智能体反馈机制依赖于截图、LLM作为评判者的评分或自然语言修正,但很少有机制能像开发者那样与实时应用程序交互。我们引入了WatchPoint,一个模拟用户系统,通过针对运行中的应用程序生成并执行诊断脚本来模拟真实开发者的行为,产生结构化观察结果以指导编码模型的重试。与先前针对单文件编辑或使用不可执行指标进行评估的方法不同,我们在Web-Bench上运行,这是一个包含50个多文件Web项目、由1000个顺序依赖任务组成的基准,通过确定性端到端测试进行验证。WatchPoint恢复了其诊断任务的57.6%,一项受控用户研究证实了模拟的真实性:人类测试者实现了可比的恢复率(54.5%),提供了自动化诊断脚本可以替代顺序Web开发任务中交互式人工测试的证据。我们进一步识别了一种能力差距模式,该模式决定了模拟用户反馈何时有帮助以及何时应被保留。
英文摘要
When a professional web developer's code fails a test, they do not simply re-read the stack trace. They open the application in a browser, click buttons, inspect computed styles, and run diagnostic commands to understand what went wrong. Existing feedback mechanisms for coding agents rely on screenshots, LLM-as-a-judge scoring, or natural-language corrections, but few interact with the live application the way a developer would. We introduce WatchPoint, a simulated-user system that mimics real developer behavior by generating and executing diagnostic scripts against the running application, producing structured observations that guide the coding model's retry. Unlike prior approaches that target single-file edits or evaluate using non-executable metrics, we operate on Web-Bench, a benchmark of 50 multi-file web projects comprising 1,000 sequentially dependent tasks, verified by deterministic end-to-end tests. WatchPoint recovers 57.6% of the tasks it diagnoses, and a controlled user study confirms the simulation's realism: human testers achieve a comparable recovery rate (54.5%), providing evidence that automated diagnostic scripts can substitute for interactive human testing on sequential web development tasks. We further identify a pattern of capability gaps that governs when simulated-user feedback is helpful and when it should be withheld.
发表机构
- Stevens Institute of Technology(史蒂文斯理工学院)
- University of Texas at Dallas(德克萨斯大学达拉斯分校)
机构由 AI 辅助整理,请以论文原文为准。