arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

动态离散选择与逆强化学习:从人类行为推断偏好与信念

Dynamic Discrete Choice and Inverse Reinforcement Learning: Inferring Preferences and Beliefs From Human Behavior

Pranjal Rawat, John Rust

arXiv 2608.24362首次发表:更新:

AI 中文总结

本文综述动态离散选择(DDC)与逆强化学习(IRL)两个紧密关联的文献,对比二者的数学表述、估计计算方法,指出交叉融合可解决共同挑战,推动方法学进步。

AI 中文摘要

本文综述了两个联系紧密的文献,它们从不同学科传统出发研究同一基本问题:结构计量经济学中的动态离散选择(DDC)与机器学习中的逆强化学习(IRL)。两者均假设个体在形式为马尔可夫决策过程(MDP)的动态不确定环境中行动以最大化期望奖励函数,旨在从观察到的序列行为中推断决策者的偏好。尽管起源独立,两个领域已收敛到相似的数学表述。我们证明,当前IRL中流行的(软Q学习)框架与带有加性极值偏好冲击的DDC模型密切相关,二者产生相同的softmax(多项logit)选择概率和支撑经济学结构估计的平滑贝尔曼方程。我们比较了两个领域开发的估计与计算方法:DDC强调最大似然估计、条件选择概率估计量和策略迭代方法;IRL开发了可扩展替代方案,包括最大熵方法、对抗方法和使用深度神经网络扩展到高维状态空间的无模型时间差分估计量。结合时间差分学习与计量经济学经典两步法的无模型IRL估计量是弥合两个文献的有前景方向。两个领域面临共同的基础挑战:识别问题(即多个奖励函数可合理化相同观察到的行为)以及求解基础MDP的维度灾难。我们认为,交叉融合为两个领域的方法学进步提供了大量机会。

英文摘要

This article surveys two deeply connected literatures that approach the same fundamental problem from different disciplinary traditions: dynamic discrete choice (DDC) in structural econometrics and inverse reinforcement learning (IRL) in machine learning. Both seek to infer the preferences of decision makers from observed sequential behavior, assuming that individuals act to maximize an expected reward function within a dynamic, uncertain environment formalized as a Markov decision process (MDP). Despite independent origins, the two fields have converged on similar mathematical formulations. We show that the (soft Q-learning) framework now prevalent in IRL is closely related to DDC models under additive extreme value preference shocks, yielding the same softmax (multinomial logit) choice probabilities and smooth Bellman equations that underpin structural estimation in economics. We compare the estimation and computational methods developed in each field. DDC has emphasized maximum likelihood estimation, conditional choice probability estimators, and policy iteration methods. IRL has developed scalable alternatives, including maximum entropy methods, adversarial approaches, and model-free temporal difference estimators that extend to high-dimensional state spaces using deep neural networks. Model-free IRL estimators that combine temporal difference learning with classical two-step methods from econometrics represent a promising direction for bridging the two literatures. Both fields confront shared foundational challenges: the identification problem, whereby multiple reward functions can rationalize the same observed behavior, and the curse of dimensionality in solving the underlying MDP. We believe that cross-fertilization offers substantial opportunities for methodological progress in both fields.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑