通过以世界模型为中心的自主智能体探索具身智能的认知-物理极限
Toward the Cognitive--Physical Limits of Embodied Intelligence through a World-Model-Centric Autonomous Racing Agent
浏览论文内容
中文总结 AI 辅助
该研究提出以世界模型为中心的自主赛车智能体,结合实车与模拟实验,探索具身智能的认知-物理极限,提升了交互成功率与泛化能力,为安全部署提供边界感知方法。
中文摘要 AI 辅助
具身人工智能旨在开发能通过与物理世界持续交互进行感知、推理和行动的智能体。然而,大多数具身系统仍在保守安全裕度或中等交互场景中评估,导致其在极端条件下的能力边界未被充分理解。自主赛车提供了严格的测试平台,它结合了高频定位与感知、对抗性交互、近饱和车辆动力学以及严格的安全约束。现有系统虽能实现高速性能,但很少联合建模和优化认知与物理极限。本文展示了一种以世界模型为中心的自主赛车智能体,为探索这些耦合极限提供了具体进展。该框架从近极限成功与失败中学习预测世界模型,以捕捉交互演化、自身动力学及可行运动边界,在闭环优化过程中耦合世界状态构建、未来感知推理与近极限控制。训练数据来自实车自主赛车,车载系统在速度达256.3 km/h、峰值横向加速度达26.8 m/s²时仍保持鲁棒的定位与感知。在全尺寸模拟赛车中,训练良好的以世界模型为中心的智能体在多种具挑战性的模拟赛车场景中达到88.3%的交互成功率。世界模型与策略的闭环优化进一步提升了对认知-物理极限的利用、故障模式的恢复能力,以及在不同条件和未见过的赛道上的泛化能力。这些结果表明,一种边界感知方法论,其中世界模型帮助具身智能体表示、预测并持续优化其能力边界,以实现更安全的现实世界部署。
英文摘要
Embodied artificial intelligence aims to develop agents that perceive, reason, and act through continuous interaction with the physical world. However, most embodied systems are still evaluated within conservative safety margins or moderate interaction regimes, leaving their capability boundaries under extreme conditions insufficiently understood. Autonomous racing provides a stringent testbed by combining high-frequency localization and perception, adversarial interaction, near-saturated vehicle dynamics, and strict safety constraints. Existing systems push high-speed performance but rarely model and refine cognitive and physical limits jointly. Here we show that a world-model-centric autonomous racing agent provides a concrete step toward exploring these coupled limits. The framework learns predictive world models from near-limit successes and failures to capture interaction evolution, ego dynamics, and feasible-motion boundaries, coupling world-state construction, future-aware reasoning, and near-limit control in a closed-loop refinement process. Training data were collected from real-vehicle autonomous racing, where the onboard system maintained robust localization and perception at speeds up to 256.3 km/h and peak lateral acceleration of 26.8 m/s$^2$. In full-scale simulated racing, the well trained world-model-centric agent achieves an 88.3% interaction success rate across various challenging simulated racing scenarios. Closed-loop refinement of the world model and policy further improved utilization of cognitive-physical limits, recovery from failure modes, and generalization across varying conditions and unseen circuits. These results suggest a boundary-aware methodology in which world models help embodied agents represent, predict, and continually refine their capability boundaries for safer real-world deployment.
发表机构
- K2 Holding, L.L.C(K2控股有限责任公司)
机构由 AI 辅助整理,请以论文原文为准。