arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

耦合精馏基准上的安全门控智能体监督控制:工况图、可审计门与协同设计发现

Safety-Gated Agentic Supervisory Control on a Coupled Distillation Benchmark: Regime Map, Auditable Gate, and Co-Design Findings

Christian Rosenthal

arXiv 2607.27849首次发表:更新:

AI 中文总结

该研究针对耦合精馏基准,提出基于规则的分叉孪生反事实安全门控机制,结合DeepSeek-V4-Flash等模型,在目标获取、扰动抑制等任务上优于基准方法,可有效遏制有害操作。

AI 中文摘要

开源权重大语言模型(LLM)可每五分钟写入设定点,但工厂仍需一项硬性检查:在控制层执行前完成命名约束、记录裕度及“允许/阻止”决策。本文将该检查置于基于规则的分叉孪生反事实门(含9个固定约束)中,控制层保持不变。在Skogestad Column A上,实验设置为仅含PID的阶梯控制(C0)、线性模型预测控制(MPC,C1)、无门控智能体(C2)及门控智能体(C3),所有方案采用相同的液位控制(M_D、M_B)、场景及随机种子;C2与C3共享线性MPC后端。对比结果显著:非标称目标获取任务中,智能体在强工况带击败帕累托调优的线性MPC(C2/C1的积分绝对误差(IAE)比率在置信区间上限为0.361);同一16点网格的扰动抑制性能在置信区间上限反转16.03(点估计为10.18),无门控LLM监督器无法达到该水平。该门将规范放弃吸引子压缩为有界偏移(d≈-1.4;95百分位单元IAE为11.5至0.77),一行提示可从源头上消除吸引子(成功率从6/10提升至0/10;仅为敏感性测试,非核心发现)。在250单元统计测试中,590次门干预里有534次为规范边界几何相关操作:运行规范处于安全极限,导致正常操作点(OP)不可用,仅能遏制异常操作;318次阻止操作仍能主动修正有害提案。核心结论基于DeepSeek-V4-Flash模型的单栏结果,NVIDIA Nemotron-3-Super的第二组测试保留了扰动抑制失效带及工厂侧故障分布;幅度与协议可操作性仍具模型依赖性,Super目标获取强工况仅为幸存者验证(非确认);本文中迁移指孪生结构、约束包络及设定点接口,未涉及第二种工厂类别的测量。

英文摘要

An open-weight LLM can write composition setpoints every five minutes. What a plant still needs is a hard check: named constraints, logged margins, and an admit/block decision before the regulatory layer moves. This paper puts that check in a rule-based forked-twin counterfactual gate (nine pinned constraints) and leaves the regulatory layer unchanged. On Skogestad's Column A the ladder is PID-only (C0), linear MPC (C1), ungated agent (C2), and gated agent (C3) under one contract: identical level closure (M_D, M_B), scenarios, and seeds; C2/C3 share the linear-MPC backend. The split is not subtle. Off-nominal target acquisition: the agent beats Pareto-tuned linear MPC in the strong band (C2/C1 IAE ratio 0.361 at the upper CI). Disturbance rejection on the same 16-point grid inverts by 16.03 at the upper CI (10.18 at the point estimate), where an ungated LLM supervisor does not belong. The gate compresses a specification-abandonment attractor into a bounded offset (d approx. -1.4; P95 cell IAE 11.5 to 0.77). A one-line prompt fix removes the attractor at source (6/10 to 0/10; sensitivity only, not a new headline). In a 250-cell statistical pass, 534 of 590 gate interventions are spec-on-bound geometry: the operating specification sits on a safety limit, so a well-behaved OP becomes inoperable while misbehaving ones are only contained; 318 blocks still correct actively harmful proposals. Headlines are single-column and model-conditional on DeepSeek-V4-Flash. A second-family sweep (NVIDIA Nemotron-3-Super) keeps the disturbance-rejection fails band and plant-side failure geography; magnitudes and protocol operability stay model-conditional, and Super target-acquisition strong cells are survivors only (not confirmation). Transfer means twin, constraint envelope, and setpoint interface, not a second plant class measured here.

Comments31 pages, 8 figures. Code and data: https://github.com/cgncro-cyber/IndustrialAI. Sole author; independent research

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑