arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.20963cs.CRcs.AI

Vibe编码与Web应用安全:一项双提示研究

Vibe Coding and Web Application Security: A Twin-Prompt Study

  • University of Zagreb(萨格勒布大学)
  • Faculty of Organization and Informatics(组织与信息学院)

机构由 AI 辅助整理,请以论文原文为准。

Darko Andročec

AI总结:

本研究通过双提示实验对比发现,明确要求安全最佳实践的Web应用生成结果,其安全问题数量更少且无高危问题,是关于LLM生成Web应用安全的初步研究。

AI中文摘要:

大型语言模型越来越多地根据自然语言提示生成完整的Web应用程序,这引发了一个问题:明确要求遵循安全最佳实践是否能提升生成结果。我们研究了六个功能不同的Web应用程序,每个程序都以两种提示变体生成,除了附加的安全要求部分外其余完全相同:基线变体(A)和安全感知变体(B)。所有十二个程序均由同一个智能体编码助手和同一模型版本在单次非迭代生成轮次中生成,之后通过静态、依赖、动态及手动技术进行分析,从85个候选问题中得出75个已确认的发现。安全感知变体在每个应用程序中产生的已确认发现更少(24个对比51个),且未包含严重或高危问题;最严重的发现仅通过手动测试检测到。由于语料库规模较小且每个变体仅生成一次,我们报告描述性观察而非统计确定的效应,并将本研究定位为初步研究,其流程正被扩展至多模型及重复运行。

英文摘要:

Large language models increasingly generate complete web applications from natural-language prompts, raising the question of whether explicitly requesting security best practice improves the result. We study six functionally distinct web applications, each generated in two prompt variants that are identical except for an appended security-requirements section: a baseline (A) and a security-aware (B) variant. All twelve programs were produced by the same agentic coding assistant and the same model version in a single, non-iterative generation round, and were then analyzed with static, dependency, dynamic and manual techniques, yielding 75 confirmed findings out of 85 candidates. The security-aware variant produced fewer confirmed findings in every application (24 versus 51) and contained no Critical or High issues; the most severe finding was detected only by manual testing. Because the corpus is small and each variant was generated once, we report descriptive observations rather than statistically established effects, and position the work as a preliminary study whose pipeline is being scaled to multiple models and repeated runs.

补充信息

↑