发表机构
Princeton University(普林斯顿大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文针对信息不对称下的人机辅助博弈,提出实用-教学最佳响应方法,可单轮解决目标不确定性,克服逆最优控制的推理上限,通过协作积木搭建示例验证了方法有效性。
AI 中文摘要
辅助博弈形式化了信息不对称下的人机协作:人类知晓目标,机器人需通过观察与交互推断目标以提供有效辅助。一般而言,在线计算最优辅助博弈策略是难解的,因为精确解需要在部分可观测马尔可夫决策过程(POMDP)中规划。本文确定了一类辅助博弈,其中实用-教学推理可在单个时间步内解决目标不确定性,使得全时域博弈可通过易处理的最佳响应程序精确求解。在该类博弈中,我们表明主流逆最优控制存在推理上限,阻碍了对齐,而实用-教学推理通过仅在任务执行下看似等价的动作立即消除目标歧义,克服了这一障碍。最后,我们在简单的协作积木搭建示例上验证了理论结果与所提方法。
英文摘要
Assistance games formalize human-robot collaboration under asymmetric information: the human knows the goal, while the robot must infer it from observation and interaction in order to assist effectively. In general, computing optimal assistance game strategies online is intractable, since exact solutions require planning in a POMDP. We identify a class of assistance games in which pragmatic-pedagogic reasoning resolves goal uncertainty in a single time step, rendering the full-horizon game exactly solvable by a tractable best-response procedure. Within this class, we show that mainstream inverse optimal control exhibits an inference ceiling that hinders alignment, while pragmatic-pedagogic reasoning overcomes this barrier by immediately disambiguating goals through actions that look equivalent under task execution alone. Finally, we validate our theoretical results and proposed method on a simple collaborative block-building example.
CommentsWorld Symposium on the Algorithmic Foundations of Robotics (WAFR) 2026