arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

同时进行的前向与逆向人在回路优化

Simultaneous Forward and Inverse Human-in-the-Loop Optimization

Kyeongwon Park, Steven H. Collins

arXiv 2609.22630首次发表:更新:

发表机构

Stanford University(斯坦福大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出SFIHILO方法,通过前向模型预测用户结果并主动查询偏好,高效逆向学习奖励函数并优化控制策略,在仿真中显著提升样本效率,实现更有效的人机交互。

AI 中文摘要

主观用户体验对人机交互至关重要,但用户重视的结果以及这些偏好如何在个体和情境之间变化往往是未知的。虽然使用人类数据的逆向学习方法有助于识别用户奖励,但在许多辅助性场景中,执行控制策略、测量生物力学或生理结果以及收集用户反馈的实验成本常常限制了查询次数和优化迭代。在此,我们提出了同时进行的前向与逆向人在回路优化(SFIHILO),该方法能高效地从对人类结果的偏好中推断出个体特定的奖励函数,并确定一个最终控制策略,以最大化学习到的奖励。SFIHILO利用前向模型从控制策略预测用户结果,使用该模型进行主动查询以加速逆向奖励学习,然后在无需额外用户试验的情况下优化最终控制策略。我们根据候选策略在前向和逆向信念中的预期不确定性减少来评分,针对区域偏好边界以稳健地指导这一同步学习过程。在仿真中,我们展示了SFIHILO在用户异质性、结果维度、精度要求、噪声水平和非平稳性方面的有效性;与互信息方法相比,所提出的主动查询策略在逆向学习中显著提高了样本效率,同时保持了前向模型的准确性。该方法展示了推断潜在人类目标的潜力,从而实现更具可迁移性和有效的人机交互。

英文摘要

Subjective user experience is important to human-robot interaction, but the outcomes users value, and how those preferences vary across individuals and contexts, are often unknown. While inverse learning approaches using human data can help identify user rewards, in many assistive settings the experimental costs of executing a control policy, measuring biomechanical or physiological outcomes, and collecting user feedback often limit the number of queries and optimization iterations. Here, we present Simultaneous Forward and Inverse Human-In-the-Loop Optimization (SFIHILO), which efficiently infers individual-specific reward functions from preferences over human outcomes and identifies a final control policy that maximizes the learned reward. SFIHILO bootstraps a forward model to predict user outcomes from control policies, uses this model for active querying to accelerate inverse reward learning, and then optimizes the final control policy without additional user trials. We score candidate policies by their expected reduction in uncertainty across both forward and inverse beliefs, targeting a regional preference boundary to robustly inform this simultaneous learning process. In simulation, we show that SFIHILO was effective across user heterogeneity, outcome dimensionalities, precision requirements, noise levels, and nonstationarity; compared with mutual information approaches, the proposed active querying strategy significantly improved sample efficiency in inverse learning while preserving forward model accuracy. This approach demonstrates the potential to infer latent human goals, enabling more transferable and effective human-robot interaction.

CommentsAccepted at the 10th Conference on Robot Learning (CoRL 2026)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑