AI 中文总结
该研究提出Spec2Vision框架,通过分阶段运行时保持任务规约明确性,在17项CV任务的850次运行评估中实现81/85的通过率,证实任务规约明确性对AI生成CV流水线交付的重要性。
AI 中文摘要
生成的计算机视觉代码可能可运行,但不满足下游评估器强制的任务规约要求。我们针对该差距研究了Spec2Vision,这是一个用于生成和评估基于规约的计算机视觉(CV)流水线 bundle 的实验框架,其通过分阶段运行时,在合成、筛选、测试和有界修复过程中保持任务规约的明确性。该基准评估了17项CV任务、10种可执行条件,且每个任务-条件单元重复5次,共850次主要运行。在850次主要运行评估中,Spec2Vision达到了81/85的评估器测试通过率;移除结构修复后降至55/85,移除兼容性支架后降至58/85,移除生成器预飞行后降至39/85。可执行单智能体基准向模型逐步提供更丰富的任务规约,最终达到直接源规约暴露,但整体仍弱得多,从轻量任务基础的17/85评估器测试通过率到35/85评估器测试通过率不等。不过,轻量基准在85/85的运行中保持核心可运行,但仅达到17/85的评估器测试通过率和6/85的严格交付成功率,表明可运行性不等同于交付。在该基准中,最有力的证据来自在分阶段生成、检查和修复过程中保持任务规约的明确性。提供了人工制品以支持对运行bundle、模型可见输入和衍生表格的审计。
英文摘要
Generated computer-vision code can be runnable without satisfying the task contract enforced by a downstream evaluator. We study that gap with Spec2Vision, an experimental framework for producing and evaluating specification-grounded CV pipeline bundles through a staged runtime that keeps the task contract explicit across synthesis, screening, testing, and bounded repair. The benchmark evaluates 17 CV tasks, 10 executable conditions, and 5 repeats per task-condition cell, for 850 primary runs. In the primary 850-run evaluation, Spec2Vision reaches 81/85 evaluator-test passes; removing structural repair drops to 55/85, compatibility scaffolding to 58/85, and generator preflight to 39/85. The executable single-agent baselines expose progressively richer task specifications to the model, culminating in direct source-spec exposure, yet remain much weaker overall, from 17/85 for lightweight task grounding to 35/85 evaluator-test passes. The lightweight baseline nevertheless remains core-runnable in 85/85 runs but reaches only 17/85 evaluator-test passes and 6/85 strict-delivery successes, showing that runnability is not equivalent to delivery. Across this benchmark, the strongest evidence comes from keeping the task contract explicit across staged generation, checking, and repair. Artifacts are provided to support audit of run bundles, model-visible inputs, and derived tables.
Journal refCEUR Workshop Proceedings, Vol. 4238, Proceedings of the 2nd Generative Code Intelligence Workshop (GeCoIn 2026), 2026