arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

反事实商模型:学习动作会改变什么,而非世界会发生什么

Counterfactual Quotient Models: Learning What Actions Change, Not What the World Does

Junlin Chen, Ruijie Wang, Jianxin Li

arXiv 2608.22092首次发表:更新:

发表机构

School of Computer Science and Engineering, Beihang University(北京航空航天大学计算机科学与工程学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出反事实商模型,通过移除动作间共享的未来分量,直接从同步反事实回滚学习动作依赖效应,在物理环境实验中验证其可抑制无关变异、提升动作排序效果。

AI 中文摘要

强化学习模型通常会预测完整的未来状态、观测值或特征占用情况,即便动作选择仅依赖于候选动作结果之间的差异。因此,这些模型可能会将大量统计和表征能力投入到与智能体当前选择无关的高维现象中。我们提出反事实商模型,该模型将仅在跨动作共享的分量上存在差异的动作条件未来视为等价。其规范中心表示会移除该共享分量,同时保留由建模奖励族可表达的每一对动作比较。所实现的模型从同步反事实回滚中直接学习这些动作依赖效应,使得共享随机动力学在函数逼近前抵消,而非在完整未来被预测后抵消。我们确立了所得表示的决策充分性、可识别性、共模不变性、逼近行为及后悔特性。在基于物理的环境中开展的受控实验为这些特性提供了初步证据:直接效应学习抑制了动作独立变异,支持此前未见过的奖励查询,并相对于训练以预测绝对未来的模型提升了动作排序效果。

英文摘要

Reinforcement-learning models commonly predict complete future states, observations, or feature occupancies, even though action selection depends only on differences between the consequences of candidate actions. As a result, these models may devote substantial statistical and representational capacity to high-dimensional phenomena that evolve independently of the agent's current choice. We introduce the Counterfactual Quotient Model, which treats action-conditioned futures as equivalent when they differ only by a component shared across actions. Its canonical centered representation removes this common component while preserving every pairwise action comparison expressible by the modeled reward family. The implemented model learns these action-dependent effects directly from synchronized counterfactual rollouts, so shared stochastic dynamics cancel before function approximation rather than after complete futures have been predicted. We establish the decision sufficiency, identifiability, common-mode invariance, approximation behavior, and regret properties of the resulting representation. Controlled experiments in physics-based environments provide initial evidence for these properties: direct effect learning suppresses action-independent variation, supports previously unseen reward queries, and improves action ranking relative to models trained to predict absolute futures.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑