arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

可审计性并非单一属性:强化学习中的规则重叠、行为一致性与组合

Auditability Is Not One Property: Rule Overlap, Behavioural Agreement, and Composition in Reinforcement Learning

Liu Hung Ming

arXiv 2609.28581首次发表:更新:

发表机构

PARRAWA AI(PARRAWA AI)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出将强化学习策略的可审计性分解为六个可测试属性,通过规则提取与组合协议验证,发现规则重叠不保证行为一致,融合策略仅选择现有规则而非生成新技能。

AI 中文摘要

强化学习(RL)策略通常以不透明的神经检查点形式分发,而训练日志仅显示运行发生过,并未解释策略学到了什么。我们研究独立训练的策略能否通过可审计的离散行为规则进行表示和组合。我们将可审计性定义为六个可分别测试的谓词:轨迹完整性、无损编码、规则覆盖、行为一致性、组合质量以及价值模型可靠性。我们的协议使用共享的冻结符号化器、被动规则提取、仅追加的哈希约束账本、精确环境重放,以及带有显式盲点回退的离线置信度排序仲裁。结果对该描述层施加了严格限制。规则集重叠并不蕴含行为一致性:策略可能共享符号规则,但在新状态上选择接近随机匹配的动作。因此,融合策略是在现有规则中进行选择,而非生成新技能。在一个冲突主导的任务中,一个明显的融合失败被追溯到归纳/部署不匹配:从采样动作中归纳出的规则在argmax动作下进行评估,而部署一致的重新归纳则反转了仲裁顺序。一个拟合Q的广义策略改进诊断在两个环境中也失败,限制了规则融合优于基于价值的组合这一主张。一项探索性比较支持规则融合,但其比较器是事后性的,任务部分饱和,且融合策略仍低于最强的留出演员。我们贡献的是一个证据受限的审计与组合协议,而非普遍可解释性或自主技能生成的声明。未来工作必须增加时间扩展技能、跨技能接口、组合搜索以及独立的创新性审计。

英文摘要

Reinforcement-learning (RL) policies are often distributed as opaque neural checkpoints, while training logs show that a run occurred without explaining what the policy learned. We study whether independently trained policies can be represented and composed through auditable discrete behavioral rules. We define auditability as six separately testable predicates: trace integrity, lossless coding, rule coverage, behavioral agreement, composition quality, and value-model reliability. Our protocol uses a shared frozen symbolizer, passive rule extraction, an append-only hash-bound ledger, exact environment replay, and offline confidence-ranked arbitration with an explicit blind-spot fallback. The results place strict limits on this description layer. Rule-set overlap does not imply behavioral agreement: policies may share symbolic rules while choosing near-chance-matching actions on fresh states. The fused policy therefore selects among existing rules rather than generating a new skill. On a conflict-dominated task, an apparent fusion failure is traced to an induction/deployment mismatch: rules induced from sampled actions were evaluated under argmax actions, and deployment-consistent re-induction reverses the arbitration ordering. A fitted-Q generalized-policy-improvement diagnostic also fails in both environments, limiting claims that rule fusion is superior to value-based composition. One exploratory comparison favors rule fusion, but its comparator is post hoc, the task is partly saturated, and the fused policy remains below the strongest held-out actor. We contribute an evidence-bounded audit and composition protocol, not a claim of universal interpretability or autonomous skill generation. Future work must add temporally extended skills, cross-skill interfaces, composition search, and independent novelty audits.

Comments35 pages, 4 figures, 14 tables. Experimental results cover eight random seeds on CartPole-v1 and Acrobot-v1. Code, data, and audit artifacts are available in the accompanying repository

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑