发表机构
Skelf Research(Skelf Research)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出RegLLM诊断工具,通过宪法奖励、升级标签和运行时治理评估受监管智能体AI的有界自主性,揭示配置差异对指标的影响,强调需更大规模评估。
AI 中文摘要
我们提出RegLLM,一种用于受监管智能体工作流中有界自主性的诊断工具。它检测六个可信度信号:引用有效性、来源接地、模式合规、升级正确性、宪法对齐和不安全动作率。这些信号通过其监督来源加以区分:程序化验证器、任务级升级标签或AI评审分数。一个确定性的运行时监督器阻止未接地的答案并强制升级,记录干预措施。相同的领域宪法用于指导评估、训练奖励和服务护栏。任务级应升级标签使行动与弃权(不执行)的决策成为可测量的训练信号。我们在烟雾规模上展示了该工具。一次离线参考运行(n=12)在启用治理时将升级召回率从0提升至0.67,并将不安全动作率从0.33降至0.08。两个单GPU Qwen2.5-3B LoRA/DPO试点(n=8,相同种子和评估划分)显示出显著差异:名义上相同的基于RL的配置产生0.25对0.12的任务成功率和1.0对0.5的升级召回率。一个答案质量适配器在运行A中将召回率从1.0降至0.5,但在运行B中从0.5提升至1.0。一个升级感知变体在运行B中未产生可测量的变化。这些小型试点并未确立可靠的适配器效果或生产就绪性。它们的贡献在于诊断:配置差异可能压倒有界自主性指标上明显的调优效果,从而促使需要更大的评估集和重复运行。
英文摘要
We propose RegLLM, a diagnostic harness for bounded autonomy in regulated agentic workflows. It instruments six trustworthiness signals: citation validity, source grounding, schema compliance, escalation correctness, constitutional alignment, and unsafe-action rate. Signals are distinguished by their source of supervision: programmatic verifiers, task-level escalation labels, or AI-judge scores. A deterministic runtime supervisor blocks ungrounded answers and forces escalation, logging interventions. The same domain constitution informs evaluation, training rewards, and serving guardrails. Task-level should-escalate labels make the act-versus-defer decision a measurable training signal. We demonstrate the harness at smoke scale. An offline reference run (n=12) lifts escalation recall from 0 to 0.67 and reduces unsafe-action rate from 0.33 to 0.08 when governance is enabled. Two single-GPU Qwen2.5-3B LoRA/DPO pilots (n=8, same seed and evaluation split) expose substantial variation: nominally identical RL-base configurations yield task success of 0.25 versus 0.12 and escalation recall of 1.0 versus 0.5. An answer-quality adapter changes recall from 1.0 to 0.5 in Run A, but from 0.5 to 1.0 in Run B. An escalation-aware variant produces no measurable change in Run B. These small pilots do not establish reliable adapter effects or production readiness. Their contribution is diagnostic: configuration variance can overwhelm apparent tuning effects on bounded-autonomy metrics, motivating larger evaluation sets and repeated runs.
Comments9 pages; ancillary evaluation artefacts. Previously submitted to NLLP 2026