TQL: Scaling Q-Functions with Transformers by Preventing Attention Collapse
TQL: 通过防止注意力崩溃来扩展Q函数
机构 * Stanford University(斯坦福大学)
AI总结 TQL通过防止注意力崩溃,提升Transformer在强化学习中扩展价值函数的性能,实现43%的性能提升。
作者
Robotics / Machine Learning
TQL: 通过防止注意力崩溃来扩展Q函数
机构 * Stanford University(斯坦福大学)
AI总结 TQL通过防止注意力崩溃,提升Transformer在强化学习中扩展价值函数的性能,实现43%的性能提升。
Cosmos Policy: 为视觉运动控制和规划微调视频模型
机构 * NVIDIA ; Stanford University(斯坦福大学)
AI总结 Cosmos Policy通过单阶段后训练将预训练视频模型转化为高效机器人策略,实现视觉运动控制与规划的先进性能。
π₀:一种面向通用机器人控制的视觉-语言-动作流模型
机构 * Physical Intelligence
AI总结 本文提出了一种基于预训练视觉-语言模型的流匹配架构,用于通用机器人控制,通过零样本学习和微调实现多样任务的执行。
Comments See project website for videos: https://physicalintelligence.company/blog/pi0 Published in RSS 2025
RoboReward: 通用视觉-语言奖励模型用于机器人学
AI总结 RoboReward通过构建机器人奖励数据集和基准,训练视觉-语言奖励模型,验证了其在机器人学习中的有效性,并展示了改进策略学习和缩小与人工奖励差距的成果。
PolaRiS:面向通用机器人策略的可扩展真实-仿真评估
机构 * University of Washington(华盛顿大学) ; Princeton University(普林斯顿大学) ; University of California, Berkeley(加州大学伯克利分校) ; Stanford University(斯坦福大学) ; Toyota Research Institute(丰田研究中心) ; University of Southern California(南加州大学) ; Cornell University(康奈尔大学)
AI总结 PolaRiS通过神经重建和数据协同训练,实现高保真度的机器人策略真实-仿真评估,提升仿真与现实的关联性并简化环境构建。
Comments Website: https://polaris-evals.github.io/
反馈下降:通过成对比较实现开放式的文本优化
机构 * Stanford University(斯坦福大学)
AI总结 Feedback Descent通过结构化文本反馈实现开放式文本优化,优于现有方法并在分子发现中表现突出。
视觉-语言-动作模型中人类到机器人的转移现象出现
机构 * Physical Intelligence(物理智能) ; Georgia Institute of Technology(佐治亚理工学院)
AI总结 本文提出了一种简单的方法,通过预训练视觉-语言-动作模型,使模型能从人类数据中学习并转移至机器人任务,从而提升泛化性能。
后验行为克隆:为高效强化学习微调预训练BC策略
机构 * UC Berkeley(伯克利大学) ; Stanford(斯坦福大学)
AI总结 本文提出后验行为克隆策略,通过建模示范者行为的后验分布来提升强化学习微调效果。
机器人视觉泛化中的不变性协同训练
AI总结 该研究通过引入状态相似性和观测扰动不变性辅助任务,提升机器人在多样视角、光照和干扰物体下的泛化能力,实验表明其性能比现有方法提高了18%。
Comments 14 pages, 10 figures
RoboArena: 分布式现实世界中通用机器人策略的评估
AI总结 RoboArena通过分布式众包评估,实现了对通用机器人策略在现实世界中的可扩展、可靠评估,优于传统集中化方法。
Comments Website: https://robo-arena.github.io/
机构 * Physical Intelligence
机构 * Johns Hopkins University(约翰霍普金斯大学) ; NVIDIA(英伟达) ; Stanford University(斯坦福大学) ; University of Toronto(多伦多大学)
Comments 10 pages, 5 figures, 4 tables, NeurIPS 2025
机构 * Stanford University(斯坦福大学)
Comments Project page: https://jen-pan.github.io/memer/
机构 * Carnegie Mellon University(卡内基梅隆大学) ; Stanford University(斯坦福大学)
机构 * Stanford University(斯坦福大学)
机构 * Department of XXX, University of YYY, Location, Country(XXX系,YYY大学,地点,国家) ; School of ZZZ, Institute of WWW, Location, Country(ZZZ学院,WWW研究所,地点,国家)
Comments 37 pages, 12 figures. WD and KH contributed equally; LH and JHL contributed equally
Journal ref Proceedings of the 13th International Conference on Learning Representations, 2025
机构 * Stanford University(斯坦福大学) ; Cornell University(康奈尔大学) ; University of California, Berkeley(加州大学伯克利分校)
Comments 8 pages, 6 figures. IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) 2025
机构 * Stanford University(斯坦福大学)
机构 * Stanford University(斯坦福大学) ; UC Berkeley(加州大学伯克利分校)
机构 * Stanford University(斯坦福大学) ; University of California, Berkeley(加州大学伯克利分校)
Comments ICML 2025
机构 * Stanford University(斯坦福大学)
机构 * Stanford University(斯坦福大学) ; SynthLabs
Comments Published in Foundations and Trends in Machine Learning as "A Tutorial on Meta-Reinforcement Learning". For the earlier version titled "A Survey of Meta-Reinforcement Learning", see v3 in the submission history at arXiv:2301.08028v3
Journal ref Foundations and Trends in Machine Learning: Vol. 18, No. 2-3, pp 224-384 (2025)
Comments 21 pages, 25 figures. International Conference on Robotics and Automation (ICRA) 2025
机构 * Stanford University(斯坦福大学)
Comments Videos are available at https://long-context-dp.github.io
Comments Project website: https://robotics-transformer-x.github.io
机构 * Stanford University(斯坦福大学)
Comments Accepted to Robotics: Science and Systems (RSS) 2025. Project website: https://openvla-oft.github.io/
机构 * Department of Computer Science, Stanford University(计算机科学系,斯坦福大学)
Comments Project website: https://bid-robot.github.io/