arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Lingjing:面向开放城市多智能体具身任务的仿真测试平台

Lingjing: A Simulation Testbed for Multi-Agent Embodied Tasks in Open-Ended Cities

Xiaohe Li, Yiru Wang, Junhao Fan, Mingyuan Liu, Jie Huang, Kaixin Zhang, Jiahao Li, Chen Qian, Zide Fan

arXiv 2608.08045首次发表:更新:

AI 中文总结

Lingjing 是一款面向开放城市异构多智能体具身任务的仿真测试平台,支持城市多智能体协同的可复现评估与故障诊断,通过评估视觉语言模型等揭示了环境感知等瓶颈。

AI 中文摘要

城市具身智能需要异构智能体(如无人机、地面机器人、自动驾驶汽车)在动态城市中协同作业,因此模拟器为这类协同的开发与评估提供了可扩展的基础。然而,现有平台会将不同具身形式隔离开来,使其与任务设计及评估相脱节。本文提出 Lingjing,一款面向开放城市环境下异构多智能体具身智能的仿真平台。Lingjing 基于地理数据重建并渲染动态城市,同步多个物理引擎,向智能体暴露共享的物理及结构化城市状态。其类 Gym 接口支持用户自定义的 ReAct 智能体,以及单/多智能体自然语言任务,具备可配置的星形或广播通信模式与资源约束。每个回合会生成可归因的回放内容,将智能体轨迹与通信关联到关系图变化、资源消耗及基于引擎的评估,用于系统性诊断。本文在共享引擎闭环协议下,针对9项城市任务评估了12个视觉语言模型,通过受控研究进一步考察了通信、可扩展性、鲁棒性及故障溯源。结果揭示了在环境感知与长时程执行中存在的持续瓶颈,还显示了任务依赖的协同权衡及新增能力带来的收益递减,而更繁重的工作负载会进一步降低成功率。Lingjing 提供了一个统一的测试平台,支持城市多智能体具身智能领域可复现的端到端评估与系统性故障诊断。

英文摘要

Urban embodied intelligence requires coordination among heterogeneous agents (e.g., UAVs, ground robots, and autonomous vehicles) in dynamic cities. Simulators therefore provide a scalable foundation for developing and evaluating such coordination. Existing platforms nevertheless isolate different embodiments and decouple them from task design and evaluation. We present \textbf{Lingjing}, a simulation platform for heterogeneous multi-agent embodied intelligence in open-ended urban environments. Lingjing reconstructs and renders evolving cities from geographic data, synchronizes multiple physics engines, and exposes shared physical and structured urban state to agents. Its Gym-like interface supports user-defined ReAct agents and single- or multi-agent natural-language missions with configurable star or broadcast communication and resource constraints. Each episode becomes an attribution-ready replay that links agent trajectories and communication to relation-graph changes, resource consumption, and engine-based evaluations for systematic diagnosis. We evaluate twelve vision-language models on nine urban tasks under a shared engine-in-the-loop protocol. Controlled studies further examine communication, scalability, robustness, and failure provenance. Results expose persistent bottlenecks in grounding and long-horizon execution. They also show task-dependent coordination trade-offs and diminishing returns from added capacity, while heavier workloads further reduce success. Lingjing provides a unified testbed that enables reproducible end-to-end evaluation and systematic failure diagnosis in urban multi-agent embodied intelligence.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑