基于规则对齐的小语言模型和多智能体自我校正的闭环控制
Closed-Loop Control with Rule-Aligned Small Language Models and Multi-Agent Self-Correction
浏览论文内容
中文总结 AI 辅助
研究能否对紧凑型小语言模型再训练用于控制推理并嵌入验证器引导的校正循环,通过组相对策略优化对齐模型,结合多种智能体,在随机热控模拟中取得高准确率和低延迟,支持该架构用于边缘可重构自主控制。
中文摘要 AI 辅助
自主工业运行的关键一步是能够根据自然语言需求规范创建和重新配置控制策略,且无需人工重新设计。人工智能智能体生成策略并由工厂感知验证器检查候选动作是可行路径,但实际部署受推理延迟和计算占用限制。本文研究能否对紧凑型小语言模型进行再训练用于控制推理并嵌入验证器引导的校正循环。使用通过组相对策略优化对齐的Qwen2.5 - 1.5B模型,结合动作智能体、符号/数字孪生风格验证层和重新提示智能体。在随机热控模拟中,该框架平均动作对齐准确率达91.5%,平均推理延迟3.84秒,在符号重映射下保持95%的范围内速率,支持了小语言模型+验证器架构作为边缘可重构自主控制的实用路径。
英文摘要
A key step toward autonomous industrial operation is the ability to create and reconfigure control policies from natural-language requirement specifications, with minimal or no manual redesign. In this setting, policy generation by AI agents can be a credible path when paired with a plant-aware validator (e.g., a digital twin) that can check generated candidate actions before execution. However, practical deployment is constrained by inference latency and compute footprint: large cloud-based models are often too slow, opaque, or data-sensitive for edge closed-loop use. This work investigates whether a compact Small Language Model (SLM) can be retrained for control reasoning and embedded in a validator-guided correction loop. We use a Qwen2.5-1.5B model aligned via Group Relative Policy Optimization (GRPO), combined with (i) an action agent, (ii) a symbolic/digital-twin-style validation layer, and (iii) a reprompting agent that iteratively steers outputs toward valid actions. In randomized thermal-control simulations (30 experiments with 500 steps each), the framework achieves 91.5% average action-alignment accuracy (86.3%--100% across cases) at 3.84\,s mean inference latency. Under symbolic re-mapping, it maintains a 95% in-range rate, indicating robust physical regulation despite reduced token-level agreement. These results support SLM+validator architectures as a practical path toward reconfigurable autonomous control at the edge.