arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

治理记录作为监督:验证器选择的自训练用于结构化工作流修复

Governance Records as Supervision: Verifier-Selected Self-Training for Structured Workflow Repair

Jesus Salas

arXiv 2608.18324首次发表:更新:

发表机构

Microsoft(微软)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出利用机器可验证工作流的治理记录,通过验证器选择的自训练提升模型结构化工作流修复能力,在PlanBench等数据集上实现了已接纳规划数量的显著提升,验证了该方法的有效性。

AI 中文摘要

机器可验证工作流会生成治理记录,将任务契约、模型尝试、验证器决策、已接受输出及目标来源关联起来。我们测试这些记录是否可用于监督有界模型,将偶尔或昂贵的能力整合为可靠的一次性执行。在全新的、结构不相交的PlanBench重规划案例上,Qwen3-14B thinking生成了24个被独立编写的VAL验证器接纳的规划。这些规划用于训练同一检查点以实现非思考式执行,无需神谕目标或更强的教师模型。在80个未开启的案例中,经VAL接纳的规划从1个增至57个,其中56个配对增益、零退化;thinking方法达到30个。该适配器在所有案例中均符合模式,且延迟约为thinking平均延迟的1/56。单独的配对接口修复门未通过测试。匹配的消融实验固定了源案例、52个候选池、24个目标数量、模型、配方及随机种子,仅改变目标选择。在160个新案例中,基础执行、模式选择执行、模型自选择执行及VAL选择执行分别达到1、55、69、102个已接纳规划。VAL选择相比自选择的配对净增益为+33(p=0.0000019647),且在两个难度层级均有增益。因此,在该范围内,独立语义选择相较于匹配替代方案具有关键作用。补充的更强教师模型Phi分支将基础Phi-4的已接纳规划从2提升至51个,模式有效输出从35提升至80个。早期合成实验确立了可教性、累积学习、构建鲁棒性及停止边界。研究结果支持针对有界、机器可检查能力的验证器选择监督,而非任意规划、企业有效性或无限制自我改进。

英文摘要

Machine-verifiable workflows produce governance records linking a task contract, model attempt, verifier decision, accepted output, and target origin. We test whether verifier-admitted outputs can supervise a bounded model by consolidating occasional or expensive capability into reliable one-shot execution. On fresh, structure-disjoint PlanBench replanning cases, Qwen3-14B thinking produced 24 plans admitted by independently authored VAL. They trained the same checkpoint for non-thinking execution, without oracle targets or a stronger teacher. VAL acceptance rose from 1/80 to 57/80. A prospective replication held targets, model revision, recipe, and evaluation corpus fixed across eight LoRA seeds and three inference realizations per seed. Every seed produced a clear lift: adapters reached 45/80 to 70/80 against 1/80 for every matched base report; the exact seed-level sign-flip test gave p=0.0078125. Target-selection performance was less stable. An initial matched seed gave 102/160 accepted plans after VAL selection versus 69/160 after blinded model self-selection. Across eight prospective seeds, the contrast was seed-dependent, included one clear reverse seed, and did not replicate (p=0.3672). VAL also had a positive descriptive aggregate over schema-only selection but failed its preregistered seed-level reliability gate (p=0.0703). The verifier remains the admission authority; no reliable downstream capability advantage of semantic selection is established. A complementary Phi arm supports stronger-teacher distillation. Earlier synthetic studies bound teachability, cumulative learning, transfer, and stopping. The evidence supports robust consolidation of one fixed, machine-checkable capability, not arbitrary planning, enterprise validity, or unrestricted self-improvement.

Comments28 pages, 7 figures, 13 tables. v2 adds prospective eight-seed replications: the Self-24 lift replicates, while selector capability advantages do not pass seed-level reliability gates; claims and discussion revised accordingly

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑