AI 中文总结
研究在模仿学习观测不匹配及有未观测混杂因素时的问题,提出因果软Q模仿学习和因果逆软Q学习算法,结合因果调整框架与逆强化学习目标,在长视任务中显著优于现有CIL算法。
AI 中文摘要
模仿学习能利用专家示范在未知环境中学习策略,但在观测不匹配和存在未观测混杂因素时存在困难。因果模仿学习(CIL)通过顺序π-后门准则识别调整集来提供框架。现有CIL方法适用于短视、低维设置,应用于长视、高维连续控制任务时性能不佳。我们引入因果软Q模仿学习(SQIL)和因果逆软Q学习(IQ-Learn),将因果调整框架与先进逆强化学习目标结合。算法在顺序π-后门准则的有效近似产生的因果调整状态表示上运行,利用连续控制环境因果结构将全视域调整简化为固定大小滑动窗口。在一系列混杂环境中评估发现,因果SQIL和因果IQ-Learn在长视任务上显著优于先前CIL算法,有时超越专家,而无因果意识的模仿方法无法学习有意义行为。
英文摘要
Imitation learning enables learning a policy in an unknown environment with a latent reward signal using expert demonstrations, but it struggles when the imitator's and expert's observations are mismatched and unobserved confounders are present in expert demonstrations. By identifying appropriate adjustment sets via the sequential $π$-backdoor criterion, causal imitation learning (CIL) provides a framework for approximating the expert's policy from confounded data. However, existing CIL methods, Causal Behavioral Cloning (Causal BC) and Causal Generative Adversarial Imitation Learning (Causal GAIL), are designed for short-horizon, low-dimensional settings. When applied to continuous control tasks with long horizons and high-dimensional state-action spaces, these methods exhibit poor performance: Causal BC suffers from compounding errors, Causal GAIL is unstable and sample-inefficient, and sequential $π$-backdoor adjustment becomes impractical. We introduce Causal Soft Q Imitation Learning (SQIL) and Causal Inverse soft-Q Learning (IQ-Learn), two off-policy causal imitation learning algorithms that combine the causal adjustment framework with state-of-the-art inverse reinforcement learning objectives. Both algorithms operate on causally-adjusted state representations produced by an efficient approximation of the sequential $π$-backdoor criterion, exploiting the causal structure of continuous control environments to reduce the full-horizon adjustment to a fixed-size sliding window. We evaluate all methods in a suite of confounded environments and find that Causal SQIL and Causal IQ-Learn substantially outperform prior CIL algorithms on long-horizon tasks, sometimes surpassing the expert, whereas all causally unaware imitation methods fail to learn meaningful behavior.
CommentsReinforcement Learning Journal 2026 (Also RLC 2026)