发表机构
Télécom SudParis; Institut Polytechnique de Paris(南巴黎电信学院; 巴黎理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出配对精确重置评估协议,在PushT任务上验证预测接口路由器可降低决策成本,其优势限于低计算价格,不支持因果充分性等特性。
AI 中文摘要
现有自适应推理和世界动作模型系统利用廉价阶段输出或预测的未来来分配额外计算资源。我们研究一个更具体的问题:在配对精确重置的物理结果下,源自Medium的接口能否预测切换到单独冻结的Full预测器是否足够降低特定任务的决策损失,以证明顺序开销的合理性。我们的贡献是一种配对评估和审计协议,而非新的通用路由规则:所有候选动作均从相同的重置状态执行,Medium和Full作用于相同的候选集和任务,它们的配对物理损失差值定义了路由目标。在新的PushT库(V106;1600个状态、39个任务、3个检查点对)上,冻结的预测接口路由器相较于独立Medium、独立Full以及延迟优势的仅任务路由器,降低了包含开销的决策成本。我们随后前瞻性地用该任务、当前DINO特征的维度匹配投影以及全部5个候选动作,针对更强的当前状态控制,密封了第二个1600状态的PushT确认(V107),且不收取DINO编码器延迟费用。预测接口将定价的物理决策成本降低了0.002549(状态聚类95%区间[-0.002867, -0.002238];单侧95%上界-0.002286),对全部3个检查点对均产生正向效果。受控PyBullet审计独立支持复合任务-预测机制路由器。顺序路由器仍比固定策略慢,其优势仅限于低计算价格。证据支持在测试的预测接口中存在超出一个刻意偏好的当前DINO控制的增量路由信息,但不支持因果充分性、计算节省、闭环价值或跨系列通用性。
英文摘要
Existing adaptive-inference and world-action-model systems use cheap-stage outputs or predicted futures to allocate additional computation. We study a narrower question: under paired exact-reset physical outcomes, can a Medium-derived interface predict when switching to a separately frozen Full predictor improves task-specific decision loss enough to justify sequential overhead? Our contribution is a paired evaluation and audit protocol, not a new generic routing rule: all candidate actions are executed from the same reset state, Medium and Full act on the same candidate set and task, and their paired physical-loss difference defines the routing target. On a fresh PushT bank (V106; 1,600 states, 39 tasks, three checkpoint pairs), a frozen prediction-interface router lowers overhead-inclusive decision cost relative to standalone Medium, standalone Full, and a latency-advantaged task-only router. We then prospectively seal a second 1,600-state PushT confirmation (V107) against a stronger current-state control using the task, a dimension-matched projection of current DINO features, and all five candidate actions, with no DINO encoder latency charged. The prediction interface lowers priced physical decision cost by 0.002549 (state-clustered 95% interval [-0.002867, -0.002238]; one-sided 95% upper bound -0.002286), with negative effects for all three checkpoint pairs. A controlled-PyBullet audit independently supports a composite task-prediction-regime router. The sequential router remains slower than fixed policies, and its advantage is restricted to low compute prices. The evidence supports incremental routing information in the tested prediction interface beyond one deliberately favoured current-DINO control, but not causal sufficiency, compute saving, closed-loop value, or cross-family generality.
Comments20 pages, 6 figures, 6 tables