发表机构
Columbia University(哥伦比亚大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
FlowCheck是一种约束语言,可通过界面指定氛围编码Web应用的信息流,经评估其能准确标记注入的30个约束违规,优于前沿大语言模型,可帮助氛围编码人员以界面形式表达意图并确定性检查代码。
AI 中文摘要
氛围编码的应用常存在隐性行为故障,即界面看似正常可用,但用户可见的信息并未流向预期状态或输出。我们提出FlowCheck,一种约束语言,可直接通过应用界面指定这些用户可见的信息流,约束也可在不读取代码的情况下显示和检查,且结构足够清晰,支持可靠的大语言模型(LLM)生成。FlowCheck将约束转换为确定性CodeQL分析,我们通过Claude Code生成的4个应用对其进行评估,并将其与3个编码模型作为找bug基准进行比较。结果显示,FlowCheck正确转换并标记了我们注入的全部30个约束违规,无假阳性。相比之下,前沿模型(Claude Opus 4.7、DeepSeek V3和Gemini Pro)在被提示查找相同代码中的bug时准确率显著更低,无一个达到完全准确率。该方法让氛围编码人员能以他们理解的界面形式表达意图,并针对他们不理解的代码进行确定性检查。
英文摘要
Vibe-coded applications often contain silent behavioral failures in which the interface appears functional even though user-visible information does not flow to the expected state or output. We introduce FlowCheck, a constraint language to specify these user-visible information flows directly through the application interface, where constraints can also be displayed and inspected without reading code, and are structured enough for reliable LLM generation. FlowCheck translates the constraints into deterministic CodeQL analyses, and we evaluate it across four applications generated via Claude Code, and compare with three coding models as bug-finding baselines. We find that FlowCheck correctly translates and flags all 30 of our injected constraint violations with no false positives. In contrast, frontier models (Claude Opus 4.7, DeepSeek V3, and Gemini Pro) showed significantly lower accuracy when prompted to find bugs in the same code, with none achieving full accuracy. This approach lets vibe coders state intent in terms of the interface they understand, and checks it deterministically against the code they do not.
CommentsAccepted at the 2nd ACM SIGPLAN International Workshop on Language Models and Programming Languages (LMPL '26), co-located with SPLASH/ISSTA 2026. 14 pages, 6 figures (10 pages main text, plus references and appendix)