arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.15483cs.RO

随机结果的风险感知偏好学习

Risk-Aware Preference Learning for Stochastic Outcomes

Yi-Shiuan Tung, Yuni Wu, Wei Jiang, Alessandro Roncone, Bradley Hayes

首次发表
浏览论文内容

中文总结 AI 辅助

研究在人机交互中从人类偏好学习奖励函数的问题,通过在Bradley-Terry框架内比较期望效用与累积前景理论,发现基于累积前景理论的学习者在恢复奖励函数时遗憾更低,凸显建模人类风险敏感性的重要性。

中文摘要 AI 辅助

从人类偏好中学习奖励函数是使机器人行为符合人机交互中用户期望的常用方法。多数现有方法假定人类用期望效用(EU)评估不确定结果,线性聚合结果效用与概率。但行为证据表明人类有系统的风险敏感性,对罕见负面事件过度重视并表现出损失厌恶。我们研究社会机器人导航中这种不匹配的后果,其中安全关键结果(如碰撞)罕见但后果严重。我们在Bradley-Terry偏好学习框架内将EU与累积前景理论(CPT,一种人类决策的非线性模型)进行比较。初步实验表明,当偏好由风险敏感用户生成时,基于CPT的学习者与基于EU的学习者相比,恢复奖励函数时的遗憾显著更低。我们的结果凸显了在从对随机机器人结果的偏好中学习奖励时对人类风险敏感性建模的重要性。

英文摘要

Learning reward functions from human preferences is a widely used approach for aligning robot behavior with user expectations in human-robot interaction. Most existing approaches assume that humans evaluate uncertain outcomes using expected utility (EU), aggregating outcome utilities linearly with their probabilities. However, behavioral evidence shows that humans are systematically risk-sensitive, overweighting rare negative events and exhibiting loss aversion. We study the consequences of this mismatch in social robot navigation, where safety-critical outcomes (e.g., collisions) are rare but highly consequential. We compare EU with Cumulative Prospect Theory (CPT), a nonlinear model of human decision-making, within a Bradley-Terry preference learning framework. Our preliminary experiments show that when preferences are generated by risk-sensitive users, CPT-based learners recover reward functions with substantially lower regret compared to EU-based learners. Our results highlight the importance of modeling human risk sensitivity when learning rewards from preferences over stochastic robot outcomes.

发表机构

  • University of Colorado Boulder(科罗拉多大学博尔德分校)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑