arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.04981cs.SEcs.AI

TeleGen:通过运行时遥测改进基于LLM的Web应用生成

TeleGen: Improving LLM-Based Web Application Generation via Runtime Telemetry

Yujia Luo, Haonan Zhang, Jiasi Shen, Zishuo Ding, Weiyi Shang

首次发表
浏览论文内容

中文总结 AI 辅助

TeleGen通过运行时遥测增强LLM生成Web应用的修复过程,将WebGen-Bench任务成功率提升8.5个百分点,并有效诊断隐藏执行路径的交互失败。

中文摘要 AI 辅助

大型语言模型能够根据自然语言需求生成可运行的Web应用,但许多生成的应用仍然无法通过交互式任务。现有的生成-执行-修复流水线会执行生成的应用,并利用任务结果或错误信息来指导代码修订。然而,这种反馈往往遗漏了浏览器操作与最终任务结果之间的运行时行为,使得交互级别的失败难以诊断。因此,我们提出了TeleGen,一个面向基于LLM的Web应用生成的可观测性增强框架。TeleGen对生成的应用进行插桩,在任务执行期间收集运行时遥测数据,并将原始遥测日志压缩为简洁的简报以用于修复。我们在WebGen-Bench和Web-Bench上评估了TeleGen。在WebGen-Bench上,TeleGen将任务成功率从无遥测修复的67.7%提升至76.2%,提高了8.5个百分点。在Web-Bench上,它将累积Pass@2从21.7%提升至29.8%。消融实验结果表明,运行时遥测提供了有用的诊断信号,而遥测简报使该信号更有效且使用成本更低。进一步分析表明,遥测对于涉及隐藏执行路径的失败尤其有帮助,例如导航、表单工作流以及前后端协调。

英文摘要

Large language models can generate runnable web applications from natural-language requirements, but many generated applications still fail interactive tasks. Existing generate-execute-repair pipelines execute the generated application and use task outcomes or error messages to guide code revision. However, this feedback often misses the runtime behavior between a browser action and the final task outcome, making interaction-level failures difficult to diagnose. Therefore, we propose TeleGen, an observability-enhanced framework for LLM-based web application generation. TeleGen instruments generated applications, collects runtime telemetry during task execution, and compresses raw telemetry logs into concise briefs for repair. We evaluate TeleGen on WebGen-Bench and Web-Bench. On WebGen-Bench, TeleGen improves task success from 67.7% with repair without telemetry to 76.2%, an increase of 8.5 percentage points. On Web-Bench, it improves cumulative Pass@2 from 21.7% to 29.8%. Ablation results show that runtime telemetry provides a useful diagnostic signal, while telemetry briefs make this signal more effective and less costly to use. Further analysis shows that telemetry is especially helpful for failures involving hidden execution paths, such as navigation, form workflows, and frontend-backend coordination.

发表机构

  • The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
  • University of Waterloo(滑铁卢大学)
  • The Hong Kong University of Science and Technology(香港科技大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑