发表机构
PythaLab, Yıldız Technical University, Istanbul, Turkey(PythaLab,Yıldız技术大学,伊斯坦布尔,土耳其)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究冻结小代码模型中错误条件自修复,引入PoPE方法,通过提示和权重通道对其评估,结果表明在提示通道内容消除形式安慰剂下解锁单元数有差异,权重通道各情况有不同表现,未确认内容归因优越性,PoPE是可重新测试的测量标准。
AI 中文摘要
冻结的小代码语言模型在本地部署,但在自我修复文献中,引导失败尝试后重试的信息仍未在无安慰剂对照的情况下进行测量。我们将失败的程序视为一个猜想,将执行反例视为相对于预言机的反驳,并引入了PoPE(波普尔式安慰剂对照评估):一种测量能否由同一模型实际使用证伪大语言模型生成代码的证据的方法。在PoPE中,错误内容与特定通道的安慰剂配对,在消除与任务相关的内容或打乱任务错误分配时保持预先声明的框架。冻结的小代码模型(0.5 - 1.5B)通过提示通道和权重通道(小数据适配器训练)在预注册规则下进行评估,每个臂 - 单元对有四代。在提示通道中,公共层筛选在内容消除形式的安慰剂下解锁了12个单元,而在实时错误模式臂上在40单元抗性带上解锁了10个单元;结果记录为机制无效。在权重通道中,错误内容适配器与无干预基线之间观察到8 - 8平局(p = 1.0),而SHA打乱安慰剂适配器以10次解锁领先;未确认内容归因的优越性。这些结果不构成等效性或非劣效性的证据。等效性未单独测试。研究结果仅限于公共层筛选终点;隐藏层确认按设计推迟。我们认为这不是作为信息的已编译批评消失,而是其在测试新猜想中的外部作用丧失:当从预言机学到的表示被写回生成状态时,测试被条件作用所取代。未声称有工作的JEPA - RL控制器。PoPE作为一种安慰剂对照、可重新测试的测量标准被提出。
英文摘要
Frozen small code LLMs are deployed locally, yet the information guiding a retry after a failed attempt is still measured without placebo controls in the self-repair literature. We treat a failed program as a conjecture and an execution counterexample as an oracle-relative refutation, and introduce PoPE (Popperian Placebo-controlled Evaluation): a methodology for measuring whether evidence that falsifies LLM-generated code can be used operationally by that same model. In PoPE, error content is paired with channel-specific placebos that keep the predeclared scaffold while ablating task-relevant content or deranging the task-error assignment. Frozen small code models (0.5-1.5B) are evaluated under preregistered rules through a prompt channel and a weight channel (small-data adapter training), with four generations per arm-unit pair. In the prompt channel, public-tier screening unlocked 12 units under the content-ablated form placebo versus 10 under the live error-pattern arm on a 40-unit resistant band; the result was recorded as mechanism-null. In the weight channel, an 8-8 tie was observed between the error-content adapter and the intervention-free baseline (p=1.0), while the SHA-deranged placebo adapter stayed ahead with 10 unlocks; content-attributable superiority was not confirmed. These results do not constitute evidence of equivalence or non-inferiority. Equivalence was not tested separately. Findings are restricted to the public-tier screening endpoint; hidden-tier confirmation was deferred by design. We read this not as compiled criticism disappearing as information, but as the loss of its external role in testing a new conjecture: when a representation learned from the oracle is written back into the generation state, testing is replaced by conditioning. No working JEPA-RL controller is claimed. PoPE is presented as a placebo-controlled, retestable measurement standard.
Comments54 pages, 6 figures, 17 tables. Preregistered, placebo-controlled evaluation