发表机构
The University of Tokyo; University of Illinois Urbana-Champaign; Chiba University; Tongji University; Shanghai Research Institute for Intelligent Autonomous Systems(东京大学; 伊利诺伊大学厄巴纳-香槟分校; 千叶大学; 同济大学; 上海智能自主系统研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究构建了CoDyControlBench基准,评估了6种LLMs的控制器设计性能,开发的15亿参数推理蒸馏模型可实现边缘部署,在机械臂试验中成功完成目标跟踪。
AI 中文摘要
尽管大语言模型(LLMs)在多个科学领域展现出卓越能力,但反馈控制器设计领域的探索仍较欠缺。现有基准主要聚焦于线性单自由度(DoF)系统及大型API托管模型,未明确其在复杂控制器设计任务上的性能及边缘部署的可行性。为解决这些局限,我们提出面向大语言模型的复杂动态系统到控制基准(CoDyControlBench),包含5个评估维度下的132种系统配置:自由度数量、系统类型、耦合水平、阻尼机制及控制器类型。我们对6种最先进的LLMs进行了3次独立运行评估,其中包括3种商业模型(GPT、Gemini、Claude)和3种开源模型(GLM、DeepSeek、Qwen)。GPT的设计成功率最高,达94.8%,而Qwen最低,为50.0%。在基准维度中,自由度和控制器类型表现出最大的模型平均设计成功率差异,范围分别为36.3%和17.6%,超过系统类型、耦合水平和阻尼机制对应的差异。GPT与Qwen的性能差距主要源于控制设计知识,特别是增益选择和瞬态限制机制的使用。针对边缘部署,我们通过推理蒸馏开发了一个专用的15亿参数模型。该推理蒸馏模型在CoDyControlBench上的表现优于答案蒸馏模型和基础模型,在1至6个自由度范围内保持稳定性能,并在气动人工肌肉驱动的机械臂的全部3次物理试验中实现了目标跟踪成功。这些结果建立了基准基线,并凸显了轻量级、可边缘部署的控制器设计模型的潜力。
英文摘要
Although remarkable capabilities have been demonstrated by Large Language Models (LLMs) across scientific domains, feedback controller design remains underexplored. Existing benchmarks focus mainly on linear single-Degree-of-Freedom (DoF) systems and large API-hosted models, leaving performance on complex controller-design tasks and feasibility for edge deployment unclear. To address these limitations, we introduce the Complex Dynamics-to-Control Benchmark for Large Language Models (CoDyControlBench), comprising 132 system configurations across five evaluation dimensions: number of DoF, system type, coupling level, damping regime, and controller type. Six state-of-the-art LLMs were evaluated over three independent runs, including three commercial models (GPT, Gemini, and Claude) and three open-source models (GLM, DeepSeek, and Qwen). GPT achieved the highest design success rate at 94.8\%, whereas Qwen showed the lowest rate at 50.0\%. Across the benchmark dimensions, DoF and controller type exhibited the largest model-averaged variations in design success, with success-rate ranges of 36.3\% and 17.6\%, respectively, exceeding those associated with system type, coupling level, and damping regime. Comparison of GPT and Qwen showed that their performance gap arose mainly from the control-design knowledge, particularly gain selection and the use of transient-limiting mechanisms. For edge deployment, a specialized 1.5B-parameter model was developed through reasoning distillation. The reasoning-distilled model outperformed the answer-distilled and base model on CoDyControlBench, maintained stable performance across 1-6 DoFs, and achieved successful traget tracking in all three physical trials on a pneumatic-artificial-muscle-driven robotic arm. These results establish a benchmark baseline and highlight the potential of lightweight, edge-deployable controller-design models.