arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.27743cs.SE

满足不等于解释:时间逻辑引导强化学习中的空洞性与训练影响审计

Satisfaction Is Not Explanation: Auditing Vacuity and Training Influence in Temporal-Logic-Guided Reinforcement Learning

Lorenzo Bacchiani

首次发表
浏览论文内容

中文总结 AI 辅助

针对时间逻辑引导强化学习中满足规范却可能空洞或冗余的问题,本文提出审计层,通过测量条款行使与干预有效性,区分六种状态,揭示满足背后的因果机制。

中文摘要 AI 辅助

满足其时间逻辑规范的强化学习策略只是通过了一项测试,而非一项保证论证。审阅者所关心的条款可能从未对训练产生过影响:它可能被完全规避,可能无论策略学到什么都由环境强制满足,或者相对于普通任务奖励而言是冗余的。满足概率和任务回报无法区分这些情况。本文引入了一个能够做到这一点的审计层。它衡量规范条款是否真正被行使,该作用是被强制的还是被选择的,以及测试因果关系的明显方式——弱化条款并重新训练——是否甚至有效。通常这并不有效:我们证明了在标准的基于接受度的奖励下,可比较的较弱/较强训练目标可以共享完美最优解,展示了在大量已发表规范语料库中的相关消融风险,然后表明一个恰当设计的干预能够检测其应有的效果。在标准强化学习基准和已发表的外部工件上,审计层区分了单一满足数所合并的六种状态。满足其规范的策略回答了“是否”的问题。本文追问“为何”。

英文摘要

A reinforcement learning policy that satisfies its temporal-logic specification has passed a test, not an assurance argument. The clause that matters to a reviewer may never have mattered to training: it may have been avoided entirely, forced by the environment regardless of what the policy learned, or redundant next to the ordinary task reward. Satisfaction probability and task return cannot tell any of this apart. This paper introduces an audit layer that can. It measures whether a specification clause was actually exercised, whether that role was forced or chosen, and whether the obvious way to test causation, weakening the clause and retraining, is even valid. Often it is not: we prove that comparable weaker/stronger training objectives can share perfect optima under standard acceptance-derived rewards, show related ablation hazards across a large corpus of published specifications, and then show that a properly designed intervention detects the effect it should. Across standard reinforcement learning benchmarks and published external artifacts, the audit layer separates six regimes that a single satisfaction number collapses into one. A policy that satisfies its specification has answered whether. This paper asks why.

补充信息

↑