arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.07461cs.LG

基于自动机学习的时间-因果强化学习推断

Temporal-Causal Inference for Reinforcement Learning via Automata Learning

Jan Corazza, Daniil Kaminskyi, Simon Lutz, Patrick Nossol, Hadi Partovi Aria, Zhe Xu, Daniel Neider

首次发表
浏览论文内容

中文总结 AI 辅助

提出TCIRL框架,通过自动机学习推断隐藏时间模式导致的相变原因,联合学习策略,实验验证其恢复正确DFA并匹配全信息基线。

中文摘要 AI 辅助

我们考虑在具有不可逆相变动力学的环境中的强化学习问题,该相变由隐藏的时间模式控制。智能体观察到基础状态,但无法直接观察到相变。我们将此问题形式化为一个两阶段非马尔可夫决策过程,并引入时间-因果强化学习推断(TCIRL)框架,该框架联合学习控制策略并推断相变的隐藏时间原因。TCIRL维护一个假设的确定性有限自动机(DFA)来跟踪当前活跃的相变,并通过反例驱动的基于SAT的综合来细化它。我们证明该假设几乎必然收敛到一个DFA,该DFA在所有可达标签序列上识别真实原因语言,从而为原始非马尔可夫决策过程产生最优策略。在遗传治疗网格世界和交通信号环境上的实验表明,TCIRL恢复了正确的原因DFA,并在两个领域中匹配了全信息基线。

英文摘要

We consider reinforcement learning in environments with dynamics that undergo an irreversible phase transition governed by a hidden temporal pattern. The agent observes the base state but cannot observe the phase directly. We formalize this problem as a two-phase non-Markovian decision process and introduce Temporal-Causal Inference for Reinforcement Learning (TCIRL), a framework that jointly learns a control policy and infers the hidden temporal cause of the phase transition. TCIRL maintains a hypothesis deterministic finite automaton (DFA) to track what phase is active and refines it via counterexample-driven SAT-based synthesis. We prove that the hypothesis converges almost surely to a DFA recognizing the true cause language on all attainable label sequences, yielding an optimal policy for the original non-Markovian decision process. Experiments on a genetic therapy gridworld and a traffic signal environment show that TCIRL recovers the correct cause DFA and matches the full-information baseline in both domains.

发表机构

  • Research Center Trustworthy Data Science and Security, TU Dortmund University(多特蒙德工业大学可信数据科学与安全研究中心)
  • School for Engineering of Matter, Transport, and Energy at Arizona State University(亚利桑那州立大学物质、运输与能源工程学院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑