arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

DreamLedger:机器人决策回路中用于世界模型想象的执行结算信用文件

DreamLedger: Where to Refuse World-Model Imagination Using Execution-Settled Credit

Xianyao Li, Ruitong Tian, Rui Min, Fang Xu, Eric Jing Du

arXiv 2608.23863首次发表:更新:

发表机构

University of Florida(佛罗里达大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出DreamLedger执行结算信用文件,管控机器人世界模型预测的消耗,在多模拟域和真实机械臂上验证,可减少无效想象、降低验证探测次数,适用于多模型与硬件场景。

AI 中文摘要

机器人已开始依据世界模型的预测采取行动,但目前可靠性仍仅通过瞬时的模型内部信号来体现。DreamLedger则将可靠性视为一种持久的部署对象:它是一种执行结算信用文件,记录消耗的预测在何种操作条件、区域和预测范围内得到验证,且在每次使用前都会被查阅。每一项被消耗的预测都会被注册为一项“声明”;可归因的结果会在无需标注成本的情况下,随现实的到来完成结算,归因阶段会排除受测量污染的结果,而结算监督头则补充稀疏的分箱。由此产生的信用会对消耗进行管控:低信用的预测会缩短依赖范围或触发额外观测;每次依赖事件都可通过依赖凭证和可重放日志进行审计。我们在三个模拟领域(室内飞行、桌面操作、2D导航)中对DreamLedger进行评估,将其应用于未修改的DreamerV3、TD-MPC2和V-JEPA 2-AC模型,并在真实的Franka机械臂上进行测试。在所有12个未见过的条件-范围单元中,声明失败均呈现剂量单调特性。与盲目消耗相比,信用管控规划减少了“无效想象”(即被消耗后未能兑现的声明),降幅达62%(95%置信区间为43%-81%),且成功率相当、碰撞率相近。在匹配的风险目标下,持久账本将操作中的验证探测次数从每回合1.00次降至0.36次,操作成功率为0.94,而对照组为0.98;基于结算的校准保留了适度的、与种子一致的操作点,这与原始的瞬时管控不同。该信任层可跨解码器、潜空间和令牌空间接口运行,包括基于真实机器人帧结算的V-JEPA 2-AC。在硬件上,结算在真实感知和接触噪声下仍可运行,部署失败循环会在线重新定价,且所有1062项已注册的支出均可从审计日志中重放。

英文摘要

World-model predictions inform robot actions, yet instantaneous reliability signals do not retain the outcomes of comparable past predictions. DreamLedger registers consumed predictions as claims, settles them against execution outcomes, and uses persistent execution history from comparable operating conditions, regions, and prediction horizons to estimate credit before future reliance. Replayable records connect each decision to its supporting evidence and eventual outcome. In ten-seed navigation comparisons at matched refusal volume, removing history features or resetting history increases burn rate, measured as failures per consumed prediction. An independent ten-seed manipulation replication at matched refusal volume finds that, relative to random refusal, DreamLedger lowers burn rate by 4.8 percentage points (95% CI: 0.8-8.7) and uses fewer probes. Randomized audits directly measure higher failure rates among denied candidates, and post-warmup shifts isolate the contribution of newly accumulated settlements. Franka experiments establish online deployment through replay of all 1,062 prediction uses and demonstrate a prospective gate transition: new failures lower previously high credit below a frozen threshold, triggering refusal before the next action. Task completion and verification cost characterize the trade-offs of these interventions.

Comments17 pages, 8 figures, 15 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑