arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向复杂动态系统反馈控制器设计的大语言模型基准测试与推理蒸馏

Benchmarking and Reasoning Distillation of Large Language Models for Feedback Controller Design in Complex Dynamical Systems

Zhongchao Zhou, Yixuan Xie, Wenwei Yu, Yuxi Lu, Yaonan Zhu, Qian Niu, Yutaka Matsuo, Yusuke Iwasawa

arXiv 2608.07004首次发表:更新:

发表机构

The University of Tokyo; University of Illinois Urbana-Champaign; Chiba University; Tongji University; Shanghai Research Institute for Intelligent Autonomous Systems(东京大学; 伊利诺伊大学厄巴纳-香槟分校; 千叶大学; 同济大学; 上海智能自主系统研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究构建了CoDyControlBench基准,评估了6种LLMs的控制器设计性能,开发的15亿参数推理蒸馏模型可实现边缘部署,在机械臂试验中成功完成目标跟踪。

AI 中文摘要

尽管大语言模型(LLMs)在多个科学领域展现出卓越能力,但反馈控制器设计领域的探索仍较欠缺。现有基准主要聚焦于线性单自由度(DoF)系统及大型API托管模型,未明确其在复杂控制器设计任务上的性能及边缘部署的可行性。为解决这些局限,我们提出面向大语言模型的复杂动态系统到控制基准(CoDyControlBench),包含5个评估维度下的132种系统配置:自由度数量、系统类型、耦合水平、阻尼机制及控制器类型。我们对6种最先进的LLMs进行了3次独立运行评估,其中包括3种商业模型(GPT、Gemini、Claude)和3种开源模型(GLM、DeepSeek、Qwen)。GPT的设计成功率最高,达94.8%,而Qwen最低,为50.0%。在基准维度中,自由度和控制器类型表现出最大的模型平均设计成功率差异,范围分别为36.3%和17.6%,超过系统类型、耦合水平和阻尼机制对应的差异。GPT与Qwen的性能差距主要源于控制设计知识,特别是增益选择和瞬态限制机制的使用。针对边缘部署,我们通过推理蒸馏开发了一个专用的15亿参数模型。该推理蒸馏模型在CoDyControlBench上的表现优于答案蒸馏模型和基础模型,在1至6个自由度范围内保持稳定性能,并在气动人工肌肉驱动的机械臂的全部3次物理试验中实现了目标跟踪成功。这些结果建立了基准基线,并凸显了轻量级、可边缘部署的控制器设计模型的潜力。

英文摘要

Although remarkable capabilities have been demonstrated by Large Language Models (LLMs) across scientific domains, feedback controller design remains underexplored. Existing benchmarks focus mainly on linear single-Degree-of-Freedom (DoF) systems and large API-hosted models, leaving performance on complex controller-design tasks and feasibility for edge deployment unclear. To address these limitations, we introduce the Complex Dynamics-to-Control Benchmark for Large Language Models (CoDyControlBench), comprising 132 system configurations across five evaluation dimensions: number of DoF, system type, coupling level, damping regime, and controller type. Six state-of-the-art LLMs were evaluated over three independent runs, including three commercial models (GPT, Gemini, and Claude) and three open-source models (GLM, DeepSeek, and Qwen). GPT achieved the highest design success rate at 94.8\%, whereas Qwen showed the lowest rate at 50.0\%. Across the benchmark dimensions, DoF and controller type exhibited the largest model-averaged variations in design success, with success-rate ranges of 36.3\% and 17.6\%, respectively, exceeding those associated with system type, coupling level, and damping regime. Comparison of GPT and Qwen showed that their performance gap arose mainly from the control-design knowledge, particularly gain selection and the use of transient-limiting mechanisms. For edge deployment, a specialized 1.5B-parameter model was developed through reasoning distillation. The reasoning-distilled model outperformed the answer-distilled and base model on CoDyControlBench, maintained stable performance across 1-6 DoFs, and achieved successful traget tracking in all three physical trials on a pneumatic-artificial-muscle-driven robotic arm. These results establish a benchmark baseline and highlight the potential of lightweight, edge-deployable controller-design models.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑