行为基础模型中超越最小二乘的任务推断
Task Inference Beyond Least Squares in Behavioral Foundation Models
浏览论文内容
中文总结 AI 辅助
针对行为基础模型中任务推断次优问题,提出BLS方法,平衡奖励重建误差与后续度量不匹配,理论给出上界,实验超越基线。
中文摘要 AI 辅助
行为基础模型(BFMs)旨在通过从奖励函数中推断任务向量,无需测试时策略学习即可解决广泛的下游任务。尽管高效,但由于任务向量的推断方式(通常采用普通最小二乘法,OLS),检索到的策略往往次优。OLS最小化奖励重建误差,但未约束奖励的排序,这可能导致检索到的零样本策略的后续度量偏离最优策略的后续度量。在本工作中,我们提出BLS,一种高效的测试时推断方法,平衡最小化奖励重建误差与减少后续度量不匹配。理论上,我们提供了由后续度量和奖励函数残差共同表征的次优性差距上界。实证上,我们在运动、操作和人形控制基准上对最先进的BFMs评估了BLS。BLS以可忽略的计算开销超越了现有任务推断基线。项目页面:此https URL
英文摘要
Behavioral Foundation Models (BFMs) aim to solve a wide range of downstream tasks without test-time policy learning by inferring a task vector from the reward function. While efficient, the retrieved policies are often suboptimal because of how this task vector is inferred, typically with ordinary least squares (OLS). OLS minimizes reward reconstruction error but leaves the ordering of rewards unconstrained, which can bias the successor measure of the retrieved zero-shot policy away from that of the optimal policy. In this work, we propose BLS, an efficient test-time inference method that balances minimizing reward reconstruction error with reducing successor-measure mismatch. Theoretically, we provide a suboptimality gap upper bound characterized by both successor-measure and reward-function residuals. Empirically, we evaluate BLS on top of state-of-the-art BFMs across benchmarks for locomotion, manipulation, and humanoid control. BLS outperforms existing task inference baselines with negligible computational overhead. Project page: https://embodiedai-ntu.github.io/BLS
发表机构
- National Taiwan University(国立台湾大学)
- National Yang Ming Chiao Tung University(国立阳明交通大学)
机构由 AI 辅助整理,请以论文原文为准。