arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.29175cs.HC

面向自动驾驶仿真的类人行人模型

A Human-Like Pedestrian Model for Automated Driving Simulations

  • Aalto University(阿尔托大学)
  • University of Potsdam(波茨坦大学)
  • Interdisciplinary Transformation University Austria(奥地利跨学科转型大学)

机构由 AI 辅助整理,请以论文原文为准。

Ruofeng Wang, Patrick Ebel, Philipp Wintersberger, Antti Oulasvirta

AI总结:

提出一种基于POMDP和深度强化学习的行人模型,在模拟器中训练出能复现广泛人类过街行为且可迁移、可微调的类人策略,用于自动驾驶仿真测试。

AI中文摘要:

自动驾驶车辆必须能够在各种交通场景中安全高效地与行人互动。尽管驾驶模拟器为学习此类能力提供了可扩展的测试平台,但现有的受理论启发的行人模型范围狭窄,仅限于单车道场景中的通行/不通行(go/no-go)过街决策。虽然数据驱动方法能够预测复杂情境下的行人行为,但它们在罕见的安全关键场景中缺乏足够的观测数据。在此,我们提出一种在模拟器中训练行人模型的方法,使得学习到的策略能够在真实、复杂的交通场景(包括多车道、繁忙交通和危险驾驶风格)中展现出可证明的类人行为。我们的技术贡献在于将行人-车辆交互新颖地定义为部分可观测马尔可夫决策过程(POMDP),并融入基于理论的感知、认知和运动约束。该模型考虑了人类在交通中行为的高度适应性,并模拟人们如何根据感知到的危险、时间压力和情境复杂性来调整自身反应。通过在模拟器中使用域随机化进行深度强化学习(RL)训练,该模型再现了迄今为止关于人类过街行为最广泛的实证研究结果,包括间隙接受、让行接受、犹豫和避险速度调整。我们证明了学习到的策略能够迁移到未见过的交通环境,并可通过微调进一步适应当地交通规范。综合这些结果,我们为模拟器就绪的行人模型建立了蓝图,可支持自动驾驶系统的开发与评估。

英文摘要:

Automated vehicles must be able to interact with pedestrians safely and efficiently across diverse traffic situations. Although driving simulators offer a scalable testbed for learning such capabilities, existing theory-inspired pedestrian models are narrow in scope and limited to go/no-go crossing decisions in single-lane settings. While data-driven approaches can predict pedestrian behavior in complex situations, they lack sufficient observations in rare, safety-critical scenarios. Here, we propose an approach to training pedestrian models in simulators so that learned policies generate demonstrably human-like behavior in realistic, complex traffic scenarios, including multiple lanes, heavy traffic, and dangerous driving styles. Our technical contribution is a novel definition of pedestrian-vehicle interaction as a partially observable Markov decision process (POMDP) with theory-grounded perceptual, cognitive, and motor constraints. It accounts for the highly adaptive nature of human behavior in traffic and simulates how people adjust their responses according to perceived danger, time pressure, and the complexity of the situation. When trained via deep reinforcement learning (RL) with domain randomization in a simulator, the model reproduces the broadest range of empirical findings shown so far on human crossing behavior, including gap acceptance, yielding acceptance, hesitation, and evasive speed adjustment. We show that learned policies transfer to unseen traffic environments, and can be further adapted to local traffic norms with finetuning. Together, these results establish a blueprint for simulator-ready pedestrian models that can support the development and evaluation of automated driving systems.

补充信息

↑