arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

指令重复:一种推理时的控制原语

Instruction Duplication as an Inference-Time Control Primitive

Victor Lavrenko

arXiv 2609.04024首次发表:更新:

发表机构

PeaceTech VC(PeaceTech VC)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出指令重复作为推理时控制原语,经多模型医学多项选择任务验证,可提升All-8诊断等指标,在答案工程场景中能显著改善轨迹修复效果,是低复杂度的实用控制方法。

AI 中文摘要

程序指令遵循是可控语言模型系统的基本要求,尤其在对生成的轨迹进行下游检查或修复时。我们提出指令重复(instruction duplication),这是一种极简的黑盒推理时控制方法,仅重复程序指令,无需重新训练或修改解码过程。在7个指令微调模型、300道医学多项选择题、8种放置条件及16800次计划生成任务中,将指令副本数量从1个增加到2个,使确定性All-8诊断(通过全部8项可观测测试的响应)从90.22%提升至93.17%(提升2.95个百分点),消除了1个副本后剩余失败案例的30.2%。预临时TF-IDF召回率从73.44%提升至74.81%(提升1.38个百分点;经Holm校正的p值<0.001),而最终答案准确率保持60.21%不变。过早承诺率从1.52%上升至2.30%(p_Holm=0.00536)。盲法挑战审计获得10/30的方向确认、20/30的感知平局,无反转;其预设的28/30确认标准未达标。然而,当下游系统基于生成的轨迹执行操作时,这种差异具有实际意义。在答案工程(Answer Engineering, AE)中,明确的轨迹状态决定局部修复,已发表的先不编辑的SSNHL终点为25.1%;仅系统的AE被复现为84.2%,添加相同的尾部副本后提升至97.1%。对于传导性诊断分支保留,对应数值为:已发表的不编辑值58.9%、复现AE后78.6%、AE加重复后73.8%——虽为AE内的下降,但仍比不编辑基线高14.9个百分点。因此,指令重复是一种低复杂度、放置敏感的控制方法,其实际价值可通过使用暴露轨迹的下游系统体现。

英文摘要

Procedural instruction following is a basic requirement for controllable language-model systems, especially when generated trajectories are inspected or repaired downstream. We introduce instruction duplication, a minimal black-box inference-time control that repeats only the procedural instruction, without retraining or decoding changes. Across seven instruction-tuned models, 300 medical multiple-choice questions, eight placement conditions, and 16,800 scheduled generations, moving from one to two copies raises the deterministic All-8 diagnostic--responses passing all eight observable tests--from 90.22% to 93.17% (+2.95 percentage points), eliminating 30.2% of the failures remaining after one copy. Pre-provisional TF-IDF recall rises from 73.44% to 74.81% (+1.38 points; Holm-adjusted p < .001), while final-answer accuracy remains exactly 60.21%. Premature commitment increases from 1.52% to 2.30% (p_Holm = .00536). A blinded challenge audit yields 10/30 directional confirmations, 20/30 perceptual ties, and no reversals; its prespecified 28/30 confirmation criterion is not met. Yet this distinction can matter operationally when a downstream system acts on the generated trajectory. In Answer Engineering (AE), where explicit trajectory state determines local repair, the published reason-first no-editing SSNHL endpoint was 25.1%; system-only AE was later reproduced at 84.2%, and the same trailing duplicate raised it to 97.1%. For conductive diagnostic branch preservation, the corresponding values are 58.9% published without editing, 78.6% with reproduced AE, and 73.8% with AE plus duplication--a within-AE decrease, but still 14.9 points above the no-editing baseline. Instruction duplication is therefore a low-complexity, placement-sensitive control whose practical value can emerge through the downstream system that consumes the exposed trajectory.

Comments7 pages, 2 tables. Code and frozen reproduction artifacts: https://github.com/victorlavrenko/answer-engineering/releases/tag/instruction-duplication-arxiv-v1

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑