arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AfterVibe:对话结束后留下了什么

AfterVibe: What Remains When the Conversation Ends

Matteo Paltenghi, Satish Chandra

arXiv 2607.09900首次发表:更新:

AI 中文总结

AfterVibe框架从vibe编码会话恢复自然语言规范,用语言模型提取规范并经再生测试验证,在72个项目上评估发现其恢复的规范抽象且强,优于人工描述,可迭代加强,意味着规范或成人工审查主要工件和记录来源。

AI 中文摘要

我们提出了AfterVibe,一个从vibe编码会话中恢复自然语言规范的框架。给定一个代码工件及其产生的对话轨迹,AfterVibe使用语言模型提取捕捉开发者意图的抽象自然语言规范,并通过再生测试进行验证:一个盲AI代理仅根据规范重新实现工件,生成的代码通过多层验证管道与原始代码进行评分。规范质量通过代理能否再生通过的代码来衡量;如果验证者认为实现等效,则规范被认为是强的,否则进行迭代优化。在来自公司内部编码会话的72个真实世界vibe编码项目上评估AfterVibe,我们发现其恢复的规范在设计上是抽象的——捕捉行为意图而不规定实现——但很强。多个独立再生在6分制中平均再生得分达到5.06,同时细节上保持多样,证实规范约束了做什么而不过度规定如何做。除了优于现有的人工编写描述外,规范还可以迭代加强到5.74分。一个实际意义是,在人工智能生成的代码超过传统代码审查时,规范而非代码可能成为人工审查的主要工件和记录来源。

英文摘要

We present AfterVibe, a framework that recovers natural-language specifications from a vibe coding session. Given a code artifact and the conversation trajectory that produced it, AfterVibe uses an LLM to extract an abstract natural-language specification capturing the developer's intent, and validates it through a regeneration test: a second, blind AI agent re-implements the artifact from the spec alone, and the resulting code is graded against the original through a multi-tier validation pipeline. Spec quality is thus measured by whether an agent can regenerate passing code; if the verifiers deem the implementations equivalent the spec is considered strong, otherwise it is iteratively refined. Evaluating AfterVibe on 72 real-world vibe-coded projects from a company's internal coding sessions, we find that its recovered specs are abstract by design-capturing behavioral intent without dictating implementation-yet strong. Multiple independent regenerations achieve a high mean regeneration score of 5.06 out of 6.0 while remaining diverse in their details, confirming that the spec constrains what without over-prescribing how. Besides outperforming existing human-authored descriptions, the specs can be further strengthened iteratively to a score of 5.74. A practical implication is that specifications-not code-could become the primary artifact for human review and the source of record at a time when AI-generated code is outpacing customary code review.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑