PhysBrain 1.5:从视觉语言模型到物理基础模型
PhysBrain 1.5: From Vision-Language Models to Physical Foundation Models
浏览论文内容
中文总结 AI 辅助
PhysBrain 1.5将视觉语言模型扩展为统一物理基础模型,通过离散序列联合优化实现理解、动作生成与状态预测,在28个基准上平均72.5分,达到开源最优并媲美专有模型。
中文摘要 AI 辅助
我们提出PhysBrain 1.5,这是一个用于理解物理环境、生成动作和预测未来状态的统一模型。受观察、交互和环境变化这一物理循环的启发,我们将这些能力整合到一个共同的学习框架中。从通用的视觉语言模型出发,我们将语言响应、末端执行器运动和密集视觉目标编码为离散序列,并通过自回归下一词元预测进行联合优化。预训练阶段完全从人类交互视频中获取具身监督,利用以任务为中心的片段将语义和空间上下文与恢复的运动及后续观察配对。随后,我们通过在有监督微调阶段混合人类演示、机器人轨迹和模拟经验来适配模型。在28个具身理解基准测试中,我们的8B模型平均得分达到72.5,创下新的开源最优水平,并与GPT-6-Astra和Gemini 3.6 Flash等领先专有模型表现相当。它在14个基准上取得了最佳开源结果,同时保留了通用多模态能力。除这些理解评估外,定性示例展示了该模型通过空间对齐的RGB、深度和机器人掩码输出生成末端执行器轨迹和预测未来场景的能力。
英文摘要
We present PhysBrain 1.5, a unified model for understanding physical environments, generating actions, and predicting future states. Motivated by the physical loop of observation, interaction, and environmental change, we bring these capabilities into a common learning framework. Starting from a general vision--language model, we encode language responses, end-effector motion, and dense visual targets as discrete sequences and jointly optimize them with autoregressive next-token prediction. Pre-training draws its embodied supervision entirely from human interaction videos, using task-centered episodes to pair semantic and spatial context with recovered motion and subsequent observations. We then adapt the model through supervised fine-tuning on a mixture of human demonstrations, robot trajectories, and simulated experience. Across 28 embodied understanding benchmarks, our 8B model achieves an average score of 72.5, setting a new open-source state of the art and performing on par with leading proprietary models such as GPT-6-Astra and Gemini 3.6 Flash. It achieves the best open-source results on 14 benchmarks while retaining general multimodal capabilities. Beyond these understanding evaluations, qualitative examples show the model's ability to produce end-effector trajectories and predict future scenes through spatially aligned RGB, depth, and robot-mask outputs.