学习序列出行选择:基于逆强化学习与模仿学习的路径与活动选择综述
Learning Sequential Mobility Choice: A Review of Route and Activity Choice through Inverse Reinforcement Learning and Imitation Learning
浏览论文内容
中文总结 AI 辅助
本综述构建了将交通选择建模与IRL、IL结合的统一框架,综述多种相关学习方法,提出嵌入行为约束的混合模型可提升出行选择预测性能。
中文摘要 AI 辅助
路径选择与活动选择是同一序列出行决策问题的相互关联层面:活动选择决定人们做什么、在哪里做以及何时做,而路径选择则支配人们在活动间的移动方式。本综述构建了一个统一框架,将交通选择建模与逆强化学习(IRL)、模仿学习(IL)相结合。在明确假设下,递归logit、logit动态离散选择以及最大熵逆强化学习共享一种软贝尔曼表示,而轨迹占有率与网络流满足相关守恒定律。不过,效用、奖励、策略、占有率、约束以及观测误差仍是不同的估计量,具有不同的行为和反事实解释。我们综述了约束学习与逆约束学习、占有率比方法与DICE方法、不完整及混合质量演示、图与序列学习、迁移学习、数据融合、多智能体选择以及大语言模型。我们的核心观点是,机器学习在嵌入行为约束框架时能发挥最大价值:精确的转移确保可行性,结构化奖励保留可解释的权衡,观测模型处理异构数据源,网络或均衡求解器生成一致的系统结果。此类混合模型可在不牺牲行为识别或政策相关性的前提下,提升可扩展性与预测性能。
英文摘要
Route and activity choice are distinct transportation problems that both require models of feasible decisions unfolding over networks and time. This critical integrative review connects transportation choice modeling with inverse reinforcement learning (IRL) and imitation learning (IL), while distinguishing evidence from transportation applications, transferable methods from other fields, and emerging proposals. We develop a four-layer sequential mobility choice framework comprising the environment, behavioral objective, stochastic choice mechanism, and observation process. Under stated assumptions, recursive logit, logit dynamic discrete choice, and maximum-entropy IRL use the same soft Bellman recursion linking future opportunities to current choice probabilities. Expected state-action visitation also satisfies conservation equations analogous to network flows. These mathematical connections do not make utility, reward, policy, occupancy, constraints, and observation error behaviorally interchangeable. Transportation evidence is strongest for network-scale planning, context-dependent reward learning, inference from incomplete trajectories, and activity-schedule generation, but remains limited for actual interventions and transfer across networks. We therefore propose a behaviorally disciplined hybrid architecture that keeps feasible actions, interpretable trade-offs, observation processes, and system feedback explicit while using machine learning for scalable computation, contextual representation, heterogeneity, and data integration.
发表机构
- School of Computing and Information Systems, Singapore Management University(新加坡管理大学计算与信息学院)
机构由 AI 辅助整理,请以论文原文为准。