arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.15043cs.AI

SCOPE:用于视频世界模型的分数隔离智能体优化

SCOPE: Score-Isolated Agentic Optimization for Video World Models

Yuhua Jiang, Jiaming Wang, Qingbin Liu, Feifei Gao

首次发表
浏览论文内容

中文总结 AI 辅助

本研究针对视频世界模型推理时改进的评估难题,提出SCOPE框架,通过类型化状态更新与冻结策略实现可审计适应,在Physics-IQ基准上较冻结基线提升+14.24,同时揭示推理时更新的益处需原则性部署机制。

中文摘要 AI 辅助

视频世界模型越来越多地被用作规划和具身决策的模拟器,但在推理时改进它们会带来一个微妙的评估问题:提示、采样器、验证器和选择器可能会共同演变,使得难以归因于收益或防止保留的反馈塑造最终策略。我们引入SCOPE(Score-Isolated Agentic Optimization,分数隔离智能体优化),这是一种用于冻结视频世界模型的可审计推理时适应的框架。SCOPE将外部控制表示为类型化状态,仅通过开发证据支持的有界变化来更新此状态,并在保留评估之前冻结生成的策略。在Physics-IQ基准上,SCOPE比精确冻结的基线提高了+14.24(95%置信区间[+8.10,+21.23])。受控的 ablation( ablation 指 ablation study,即消融实验)进一步确定了场景指定、采样和学习选择带来的收益,而与最强匹配的智能体基线之间的差距仍未解决。跨骨干网络和前瞻性评估揭示了一个互补结果:有用的推理时更新确实存在,但它们的益处不会在模型和设置之间均匀转移。总之,这些发现表明,可靠的推理时适应不仅需要更好的提案,还需要一种原则性机制来决定哪些更新应成为已部署系统的一部分。代码可在此URL获取。

英文摘要

Video world models are increasingly used as simulators for planning and embodied decision making, yet improving them at inference time introduces a subtle evaluation problem: prompts, samplers, verifiers, and selectors may evolve together, making it difficult to attribute gains or prevent held-out feedback from shaping the final policy. We introduce \scope (\emph{\scopefullname}), a framework for auditable inference-time adaptation of frozen video world models. \scope represents external controls as a typed state, updates this state only through bounded changes supported by development evidence, and freezes the resulting policy before held-out evaluation. On Physics-IQ benchmark, \scope improves over the exact frozen base by $+14.24$ (95\% CI $[+8.10,+21.23]$). Controlled ablations further identify gains from scene specification, sampling, and learned selection, while the margin over the strongest matched agentic baseline remains unresolved. Cross-backbone and prospective evaluations reveal a complementary result: useful inference-time updates exist, but their benefits do not transfer uniformly across models and settings. Together, these findings suggest that reliable inference-time adaptation requires not only better proposals, but also a principled mechanism for deciding which updates should become part of the deployed system. Code is available at https://github.com/YuhuaJiang2002/SCOPE.

发表机构

  • Tsinghua University(清华大学)
  • National University of Singapore(新加坡国立大学)
  • Tencent(腾讯)

机构由 AI 辅助整理,请以论文原文为准。

↑