从特权控制到可部署适应:融合机制引导的任务约简与学习行为
From Privileged Control to Deployable Adaptation:Fusing Mechanism-Guided Task Reduction with Learned Behavior
浏览论文内容
中文总结 AI 辅助
该研究针对输入增益变化与加性扰动的控制问题,提出机制引导的任务约简方法,将特权多工况控制知识转化为可部署自适应控制器,使跟踪RMSE降低约69%。
中文摘要 AI 辅助
输入增益的同时变化与大幅加性扰动构成了一类控制问题,其中固定观测器或标称控制器可能无法复现感知工况设计的性能。本文研究训练与部署间的不对称性:在仿真或调试阶段,专家控制器可利用已知的增益与扰动信息,而部署后的控制器仅能使用参考值与测量状态。直接模仿专家动作通常存在安全风险,因为同一瞬时的学生观测可能对应不同的特权工况,进而产生不同的专家动作。本文提出一种机制引导的迁移路径,而非新的神经架构:精确采样数据恒等式可消除专家控制律中的加性扰动,将学习简化为从因果状态历史推断的与任务相关的逆输入增益;潜在目标由专家动作与部署可见轨迹重构,因此无需将真实被控对象参数作为学生标签。针对实际增广采样递归推导了公共二次证书,并给出显式残差、覆盖度、切换、噪声与饱和约束。参数工况扫描显示,标称观测器的误差随a减小而急剧增大,且对于所有测试的a<1,最佳非故障调谐之上的下一个增益会发散;直接动作网络虽离线误差适中,但闭环表现仍不佳,而结构化学生控制器始终接近特权专家,在未见过的60秒试验中,跟踪RMSE较调谐观测器降低约69%。本文的贡献是提供了一种可解释的设计视角,用于将特权多工况控制知识转化为可部署的自适应控制器,并给出该迁移具有意义的条件。
英文摘要
Simultaneous input-gain variation and large additive disturbance create a control problem in which a fixed observer or nominal controller may be unable to reproduce the performance of a regime-aware design. We study a training--deployment asymmetry: during simulation or commissioning, an expert controller is allowed to use the known gain and disturbance, whereas the deployed controller can use only the reference and measured states. Directly imitating expert actions is generally unsafe because the same instantaneous student observation may correspond to different privileged regimes and hence different expert actions. We propose a mechanism-guided transfer route rather than a new neural architecture. An exact sampled-data identity removes the additive disturbance from the expert law and reduces learning to a task-relevant inverse input gain inferred from causal state history. The latent target is reconstructed from expert actions and deployment-visible trajectories, so the true plant parameter is not required as a student label. A common-quadratic certificate is derived for the actual augmented sampled recursion, followed by explicit residual, coverage, switching, noise, and saturation qualifications. A parameter-regime scan shows that the nominal observer's error grows sharply as $a$ decreases and that the next gain above the best non-failing tuning diverges for every tested $a<1$. Direct action networks also fail in closed loop despite moderate offline error, whereas the structured student remains close to the privileged expert and reduces tracking RMSE by about 69\% relative to the tuned observer in unseen 60-s trials. The contribution is an interpretable design perspective for turning privileged multi-regime control knowledge into a deployable adaptive controller, together with conditions under which the transfer is meaningful.