一种受模式约束的语言模型,用于在不破坏局部学习稳定性的情况下优化机器人策略
A Schema Bounded Language Model for Refining Robot Policies Without Destabilizing Local Learning
- Northern Arizona University(北亚利桑那大学)
- New Jersey Institute of Technology(新泽西理工学院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
该研究针对分散式系统中异构复合机器人导航问题,提出受模式约束的LLM优化机器人策略,通过跨LLM通信和多组件协同,在仿真中实现最优导航性能。
AI中文摘要:
本文研究了分散式系统中异构复合机器人的导航问题,其中策略推理与局部控制在不同更新级别运行。在NetLogo-Python实现中,三个机器人共享运动动力学但使用不同的大语言模型(LLM)后端,每个机器人独立结合LLM策略智能体、上置信界(UCB)多臂老虎机和双深度Q网络(Double DQN)控制器,无中央LLM生成团队动作。LLM推理限于回合级策略生成与优化,而非滴答级动作选择。机器人通过包含策略、结果和学习反馈的共享回合摘要进行跨LLM通信;UCB执行优化模式选择,基于策略的Double DQN从导航变量、活动策略参数和LLM动作先验中执行滴答级动作选择。对四种配置各评估30回合,固定仿真中,完整配置在全部90个相关机器人-回合记录中均到达目标,且实现最低中位数完成时间(42滴答)和P90(73.2滴答),其中位数较其他配置低25.0%至39.1%;这些观察为评估的配置提供了描述性的配置级证据。
英文摘要:
This paper addresses navigation by composite heterogeneous robots in a decentralized system when policy reasoning and local control operate at different update levels. In a NetLogo--Python implementation, three robots share motion dynamics but use different LLM backends. Each robot independently combines a large language model (LLM) policy agent, an Upper Confidence Bound (UCB) bandit, and a Double Deep Q-Network (Double DQN) controller; no central LLM generates team actions. LLM inference is confined to round-level policy generation and refinement rather than tick-level action selection. The robots perform cross-LLM communication through a shared round summary containing policies, outcomes, and learning feedback. UCB performs refinement-mode selection, and the policy-conditioned Double DQN performs tick-level action selection from navigation variables, active policy parameters, and the LLM action prior. Each of the four configurations was evaluated over 30 rounds. In the fixed simulation, the complete configuration reached the goal in all 90 correlated robot--round records and achieved the lowest median completion time (42 ticks) and P90 (73.2 ticks); its median was 25.0--39.1\% lower than those of the other configurations. These observations provide descriptive, configuration-level evidence from the evaluated configurations.