发表机构
Tampere University(坦佩雷大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出一种自适应共享控制方法,利用k级有界理性模型估计人类行为,通过分布感知最佳响应调整机器人辅助,在模拟中降低KL散度和累积成本。
AI 中文摘要
本文研究了非线性控制仿射系统的自适应共享人机控制,其中放宽了人类完全理性的假设,机器人根据观察到的有界理性人类行为调整其辅助。我们使用两人博弈的k级有界理性模型,通过交替最佳响应计算构建候选人类和机器人策略的有限库,并使用自适应动态规划近似相关的价值函数和策略。在共享控制交互过程中,状态转移残差将测量的系统演化与候选人类策略预测的轨迹进行比较。残差使用遗忘因子累积,并映射到有限候选库上的概率性人类行为模型。机器人不是选择单一候选或平均存储的机器人策略,而是通过最小化完整估计人类行为分布上的期望合作成本,计算分布感知的一步最佳响应。对于二次终端值近似和欧拉状态传播,该响应以期望人类输入的形式具有闭式解。所提出的方法在基准非线性系统稳定任务和平面机械臂共享控制设置的模拟中进行了评估。报告的结果显示,估计与模拟人类行为分布之间的Kullback-Leibler散度降低,并且在共享控制交互期间,机器人代理的累积运行成本低于最大概率和概率加权替代策略基线。
英文摘要
This work considers adaptive shared human-robot control for nonlinear control-affine systems, where the assumption of a fully rational human is relaxed and the robot adapts its assistance to observed boundedly rational human behavior. We use a level-k bounded-rationality model of the two-player game to construct a finite bank of candidate human and robot policies through alternating best-response computations, with the associated value functions and policies approximated using adaptive dynamic programming. During the shared-control interaction, state-transition residuals compare the measured system evolution with the trajectories predicted by the candidate human policies. The residuals are accumulated using a forgetting factor and mapped to a probabilistic human-behavior model over the finite candidate bank. Rather than selecting a single candidate or averaging stored robot policies, the robot computes a distribution-aware one-step best response by minimizing an expected cooperative cost over the complete estimated human behavior distribution. For a quadratic terminal-value approximation and Euler state propagation, this response admits a closed-form solution expressed in terms of the expected human input. The proposed methods are evaluated in simulations of a benchmark nonlinear system stabilization task, and of a planar manipulator shared control setup. The reported results show decreasing Kullback-Leibler divergence between the estimated and simulated human behavior distributions, and a lower accumulated running cost for the robot agent over the shared control interaction period, than the maximum-probability and probability-weighted alternative policies baseline.
Comments20 pages, 19 figures