arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

权衡 proximity 与安全:人群中跟随目标行人的约束分解方法

Navigating the Proximity-Safety Balance: Constraint Decomposition for Human Following in Pedestrian Crowds

Shiting Gong, Jianpeng Yao, Jinfeng Wang, Marco Pavone, Jiachen Li

arXiv 2608.10056首次发表:更新:

发表机构

University of Pennsylvania; University of California, Riverside; Stanford University; NVIDIA Research; Georgia Institute of Technology(宾夕法尼亚大学; 加州大学河滨分校; 斯坦福大学; 英伟达研究院; 佐治亚理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对人群中行人跟随的 proximity 与安全平衡问题,提出多约束强化学习框架,量化行人运动不确定性以增强安全性,经实验和真实机器人部署验证了方法的有效性。

AI 中文摘要

在拥挤环境中跟随目标行人时,需要在靠近目标与在周围行人和障碍物间安全穿行之间取得平衡,这种冲突在密集场景中更为突出:激进跟随易引发碰撞,保守的安全边界则会导致目标丢失,尤其是当行人行为陌生或不可预测时。现有强化学习(RL)方法通常将这些相互冲突的目标编码为单一的密集奖励,由此产生的 proximity-安全平衡是隐含的,难以跨场景调整。为解决该问题,我们将行人跟随任务分解为稀疏任务奖励和多约束强化学习框架下的独立成本约束,每个约束通过具有直接行为意义的成本阈值管理,而非隐含的奖励权重比例,从而可对权衡进行显式且可调的控制。我们进一步量化行人运动的预测不确定性,并将这些估计值整合到 RL 成本中,以在不可预测的条件下提升安全性。在分布内和分布外设置下的大量实验表明,与基线方法相比,我们的方法实现了有效的 proximity-安全平衡,真实机器人部署进一步验证了该方法在现实场景中的可行性。更多详情可访问我们的项目页面:this https URL。

英文摘要

Following a target human in crowded environments involves an inherent conflict between staying close to the target and navigating safely among surrounding pedestrians and obstacles. This conflict becomes more severe in dense scenarios, where aggressive following risks collisions and conservative margins lead to target loss, especially when pedestrian behaviors are unfamiliar or unpredictable. Existing reinforcement learning (RL) methods typically encode these competing objectives into a single dense reward, but the resulting proximity-safety balance is implicit and difficult to adjust across conditions. To address this, we decompose the human-following task into a sparse task reward and independent cost constraints within a multi-constraint RL formulation, where each constraint is managed through cost thresholds with direct behavioral meaning rather than implicit reward weight ratios, allowing explicit and tunable control over the trade-off. We further quantify the prediction uncertainty of human motions and integrate these estimates into the RL costs to enhance safety under unpredictable conditions. Extensive experiments across both in-distribution and out-of-distribution settings demonstrate that our method achieves an effective proximity-safety balance compared to baselines. Real-robot deployment further validates the feasibility of our method in real-world scenarios. More details are available on our project page: https://nav-ps-balance.github.io/.

CommentsIEEE/RSJ International Conference on Intelligent Robots and Systems (IROS 2026); Project Website: https://nav-ps-balance.github.io/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑