发表机构
LawZero(LawZero)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
PyINE框架通过插桩Python代码执行生成可验证轨迹与任务变体,用于研究推理模型在捷径行为下的监督问题,并评估多种监督方法对失败模式的覆盖能力。
AI 中文摘要
推理模型在解决任务时可能仍然保持能力,但同时默认采用更便宜但具有误导性的捷径。这产生了一个核心的监督问题:当模型给出一个带有看似合理但不完整推理的答案时,监督者能否确定该输出是否应被信任?为了研究这个问题,我们引入了PyINE,一个使用插桩的Python程序作为可验证执行基质的可扩展启发与监督框架。在PyINE中,程序定义任务环境,执行轨迹为结果和中间事实提供权威标签,任务变体可以通过机械方式生成,而非通过静态的人工标注。我们将该框架实例化为PyINE-v1,这是首个版本,由近一百万条确定性执行轨迹和超过50万个匹配的LLM生成的代码变体构建,用于反事实评估。使用带有可验证奖励的标准强化学习,在提示变体任务上,我们训练了一个遵循捷径的模型,该模型在预测执行结果方面显著改进,但当误导性的人类可见提示与程序的实际行为冲突时,仍会犯系统性错误。然后,我们评估了激活探针、训练好的文本分类器、提示判断器以及一种轻量级辩论协议作为该模型的监督者。我们发现,在数据集层面汇总的性能可能掩盖对最重要失败模式的薄弱覆盖:廉价的习得监督者经常遗漏罕见的捷径驱动错误,而更强的基于模型的检查则更为平衡,但成本显著更高,且更难转化为可靠的阈值化决策。PyINE-v1将这种失败模式覆盖问题转化为一个可复用的实验设置,用于开发可验证、对失败模式敏感且成本敏感的监督方法。
英文摘要
Reasoning models can remain capable of solving a task while still defaulting to cheaper but misleading shortcuts. This creates a central oversight problem: when a model gives an answer with plausible but incomplete reasoning, can an overseer determine whether that output should be trusted? To study this problem, we introduce PyINE, a framework for scalable elicitation and oversight using instrumented Python programs as a verifiable execution substrate. In PyINE, programs define task environments, execution traces provide authoritative labels for outcomes and intermediate facts, and task variants can be generated mechanically rather than through static human annotation. We instantiate the framework in PyINE-v1, a first release built from nearly one million deterministic execution traces and over 500,000 matched LLM-generated code variants used for counterfactual evaluation. Using standard RL with verifiable rewards on cue-varied tasks, we train a shortcut-following model that improves substantially at predicting execution outcomes while still making systematic errors when misleading human-facing cues conflict with the program's realized behavior. We then evaluate activation probes, trained text classifiers, prompted judges, and a lightweight debate protocol as overseers of this model. We find that performance pooled at the dataset level can hide weak coverage of the failures that matter most: cheap learned overseers often miss rare shortcut-driven errors, while stronger model-based checks are more balanced but substantially costlier and harder to turn into reliable thresholded decisions. PyINE-v1 turns this failure-mode coverage problem into a reusable experimental setting for developing oversight methods that are verifiable, failure-mode-aware, and cost-sensitive.