arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.09537cs.AI

VERDI:检索并非持续世界模型优化的迁移

verdi: retrieval is not transfer for continual world model optimization

Junyu Wu, Shiqin Nie, Youyi Kou, Baohua Yin, Guocai Yao, Qingyu Chen, Jingheng Ma, Shiji Zhou, Hongyong Song, Mingchen Zhuge, Sen Cui, Changshui Zhang

首次发表
浏览论文内容

中文总结 AI 辅助

本文针对基础世界模型优化中策略难迁移的问题,提出VERDI框架,通过构建优化指纹、检索假设并经目标侧验证来实现持续优化,可降低搜索与GPU成本,减少负迁移并提升迁移结果预测准确率。

中文摘要 AI 辅助

基础世界模型在规划、仿真和具身智能领域已取得显著进展,但针对用户指定目标优化预训练世界模型仍存在困难:每次优化流程通常需从头重新发现优化策略,且产生的知识极少能迁移至下一个模型。现有研究智能体可实现优化流程自动化,但将成功策略视为可直接复用的方案,未为迁移的适用场景提供原则性保障。本文提出“检索并非迁移”的观点:在一个模型上验证的策略最多只是另一个模型的优化假设,仅在目标侧实验验证后才成为可迁移知识。基于该原则,本文提出VERDI,一种基于证据许可的持续世界模型优化框架。VERDI通过共享推理时探针表征每个世界模型,构建优化指纹;检索相关先验经验作为排序后的假设;在冻结的目标侧验证器下验证每个候选方案,仅当其通过验证后才允许作为可复用证据;邻近指纹间的矛盾进一步触发探针演化,持续优化诊断表征本身。在Ctrl-World、Cosmos系列和RoboCoin上的实验表明,VERDI可降低68%的搜索成本、69%的GPU成本,将负迁移从0.34降至0.06,同时以83%的符号准确率预测迁移结果。

英文摘要

Foundation world models have made remarkable progress in planning, simulation, and embodied intelligence. However, optimizing a pretrained world model toward a user-specified objective remains difficult: each campaign typically rediscovers optimization strategies from scratch, and the resulting knowledge rarely transfers to the next model. Existing research agents automate the optimization loop but treat successful strategies as directly reusable recipes, without principled safeguards for when transfer is appropriate. We argue instead that retrieval is not transfer: a strategy validated on one model is at best an optimization hypothesis for another, and becomes transferable knowledge only after target-side experimental valida- tion. Guided by this principle, we propose VERDI , a continual framework for evidence-licensed world model optimization. VERDI characterizes each world model through shared inference-time probes to construct an Optimization Fin- gerprint, retrieves relevant prior experience as ranked hypotheses, and validates every candidate under a frozen target-side verifier before admitting it as reusable evidence; contradictions among nearby fingerprints further trigger probe evolution, continually refining the diagnostic representation itself. Experiments on Ctrl-World, the Cosmos family, and RoboCoin show that VERDI reduces search cost by 68%, GPU cost by 69%, and negative transfer from 0.34 to 0.06, while predicting transfer outcomes with 83% sign accuracy.

补充信息

↑