arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.17956cs.LGcs.AIcs.SYeess.SY

被遗漏的模式即罕见规则:连续代码世界模型中的采样-验证危险定律

An Omitted Mode Is a Rare Rule: The Sampling-Verification Danger Law in Continuous Code World Models

Javier Aguilar Martín

首次发表
浏览论文内容

中文总结 AI 辅助

该研究揭示连续代码世界模型中采样-验证机制的危险,通过理论分析与实验发现,LLM合成的模式盲模型易被利用,接受仅能证明样本一致性,无法保证连续控制任务的性能。

中文摘要 AI 辅助

在代码世界模型范式中,大语言模型(LLM)会合成一个可执行的世界模型,供经典规划器搜索,当该模型能复现采样的转移时即被接受。我们探究这种接受在连续控制中能证明什么。我们将该流程的危险定义为预期风险,并分离出其精确因子:N个独立同分布的门展开(gate rollout)均未命中概率为r的关键事件的概率恰好为(1-r)^N;一个独立的接受样本会将其预算加入指数项。在三个混合仪器上,接受的模式盲模型被利用:规划器被固定在模式边界,遗憾值接近整个可获得回报。我们证明了边界点处有效的定位预算:在某点处,利普希茨常数(Lipschitz constant)至多为L、差异为η的模型,在体积至少为κ((η-ε)/L)^(d+m)的区域上,会在容差ε以上产生分歧;所研究的不连续重置模式无需支付此类预算。在真实的LLM合成中,GPT-5.x在111次包含模式的抽取中,有105次修复了被遗漏的1维钳位(clamp)——在56个仪器流块中,每一次尝试在50个块上都精确(95%置信区间[0.781, 0.960])。在2维区域上,没有人工制品(artifact)能恢复该规则(0/156);8次针对性干预未能消除该故障,而阳性对照能定位它:无法诱导出定位规则,但若给定形式和位置,常数会完全符合。一个版本空间证书证明识别是类相关的:在最宽剂量下,宣称的拟合在20/20个块中成功,每个与样本一致的圆在18/20个块中都在容差内。我们证明了一类入口规则与每个样本完全一致,但在运行中无害,因此可识别性是仪器的可测量属性。对所有1034个人工制品在独立样本上重新评分,证实接受仅能证明样本一致性,别无其他:在门可证明具有信息性的地方,它覆盖了被利用的规划器约2%的查询。

英文摘要

In the Code World Model paradigm an LLM synthesizes an executable world model that a classical planner searches, and the model is accepted when it reproduces sampled transitions. We ask what that acceptance certifies in continuous control. We define the pipeline's danger as an expected risk and isolate its exact factor: the probability that N i.i.d. gate rollouts all miss a critical event of probability r is exactly (1-r)^N; an independent acceptance sample adds its budget to the exponent. On three hybrid instruments the accepted mode-blind model is exploited: the planner is pinned at the mode boundary at a regret of nearly the whole attainable return. We prove a localization budget, valid at boundary points: models with Lipschitz constant at most L differing by eta at a point disagree above tolerance eps on a region of volume at least kappa((eta-eps)/L)^(d+m); the discontinuous reset modes studied pay no such budget. With real LLM synthesis, GPT-5.x repairs an omitted 1D clamp in 105 of 111 mode-containing draws -- every attempt exact on 50 of 56 instrument-stream blocks (95% CI [0.781, 0.960]). On 2D regions no artifact recovers the rule (0/156); eight targeted interventions leave the failure in place, and positive controls locate it: a located rule is not induced, while given form and location the constants follow exactly. A version-space certificate proves identification is class-relative: at the widest dose the declared fit succeeds in 20/20 blocks and every sample-consistent circle is within tolerance in 18/20. We prove a class of entry rules exactly consistent with every sample yet harmless at play, so identifiability is a measurable property of the instrument. Re-scoring all 1034 artifacts on independent samples confirms acceptance certifies sample consistency and no more: where the gate is provably informative it covers about two percent of the exploited planner's queries.

发表机构

  • AGILabs

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑