AI 中文总结
本研究针对AI辅助代码生成中未明确选择被静默解决的问题,提出决策固定方法,通过枚举未指定点并构建对比实现来记录决策,在基准上实现高准确率并有效约束生成与合规检查。
AI 中文摘要
自然语言需求会留下未解决的问题,而被要求实现该需求的模型会静默地解决这些问题:在来自三个模型的600个生成测试套件中,42.7%的套件不包含能够区分竞争性解读的测试。在实现之前编写的验收示例是通常的补救措施,但一个两种解读都满足的示例无法解决任何问题。我们在这一实践中插入两个步骤:按名称枚举需求中未明确指定的点,然后针对每个点构建两个仅在这一点上有所不同的临时实现,并且仅当执行两者表明它们产生分歧时才保留候选输入。结果被记录为决策固定点:命名的点、确认的输入以及人在两个展示结果之间选择的值。同一记录随后约束生成并通过执行来决定合规性。在一个包含40个任务(配有配对参考实现)和两个模型的基准上,对于92.5%的决策点获得了区分输入,并且对于85-90%的决策点获得了识别预期决策的固定点。作为对703个独立生成实现的检查,固定点与基准的分类在96-97%的情况下一致,在三个任务上存在分歧,其中一个任务中两个分类器都出错。作为生成约束,固定点在所有210次生成中都在固定输入处得到遵守,并且在38个单元中在保留输入上从不比散文规则更差,尽管也不比陈述相同范围的散文更好。在散文规则下的每次合规失败都源于模型自行决定规则的范围;其中一个案例静默地推翻了另一个已记录的决策,通过两个模型上的留一法被归因于单一规则,并且被文本级协调遗漏,但通过重新运行记录的输入而被捕获。该设置产生的此类冲突太少,无法评估回归步骤,我们说明了原因。
英文摘要
A natural-language requirement leaves questions open, and a model asked to implement it settles them silently: across 600 generated test suites from three models, 42.7% contain no test that distinguishes the competing readings. An acceptance example written before implementing is the usual remedy, but an example both readings satisfy resolves nothing. We insert two steps into that practice: enumerate the requirement's underspecified points by name, then for each construct two throwaway implementations differing only in that point and keep a candidate input only if executing both shows they disagree. The result is recorded as a decision pin: the named point, the confirmed input, and the value the person chose between the two exhibited results. The same record then constrains generation and decides compliance by execution. On a benchmark of 40 tasks with paired reference implementations and two models, a separating input is obtained for 92.5% of decision points and a pin identifying the intended decision for 85-90%. As checks on 703 independently generated implementations, pins agree with the benchmark's classification on 96-97%, with disagreements on three tasks, one where both classifiers erred. As generation constraints, pins are honoured at the pinned input in all 210 generations and are never worse than a prose rule on held-out inputs in 38 cells, though no better than prose stating the same scope. Every compliance failure under a prose rule came from the model deciding the rule's scope itself; one such case silently overturned another recorded decision, was attributed to a single rule by leave-one-out on both models, and was missed by text-level reconciliation but caught by re-running the recorded input. The setting yields too few such conflicts to evaluate a regression step, and we say why.
Comments24 pages