发表机构
Nanyang Technological University(南洋理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出AnyviewMeter,利用相机几何与多视角注意力适配机器人奖励模型,通过低秩微调和块注意力提升单视图及多视图预测精度,显著降低误差。
AI 中文摘要
机器人奖励模型从视觉观察中评估任务执行情况,但其预测可能随相机视角和遮挡而变化,即使底层任务状态未变。因此,将预训练的奖励模型适配到本地任务需要考虑该任务是如何被观察的。我们提出AnyviewMeter,一种用于机器人奖励模型的几何条件适配框架,该模型将任务进度表示为标量奖励信号。它结合了低秩微调与令牌对齐的普吕克射线以及同步块注意力:射线条件化将相机几何融入视觉特征和注意力查询与键中,而块注意力在预训练解码器内部融合同步视图。该框架通过参数高效地适配预训练的Robometer模型,支持单视图奖励预测和联合多视图评估。在PickCube上,单视图适配改善了每个相机组的进度预测,并在视野变化下将平均绝对误差相对于RGB微调降低了约21%。在模拟操作任务中,联合多视图预测相比平均单视图RGB预测,将进度误差降低了41%-69%,并在约88%的任务-相机组中改善了时间排序。在具有固定和腕装相机的真实任务中,平均绝对误差相对于平均RGB微调降低了约21%。这些结果支持相机几何和联合视觉证据作为特定任务机器人奖励适配的有用组成部分。
英文摘要
Robotic reward models evaluate task execution from visual observations, but their predictions can change with camera viewpoint and occlusion even when the underlying task state is unchanged. Adapting a pretrained reward model to a local task therefore requires accounting for how that task is observed. We introduce AnyviewMeter, a geometry-conditioned adaptation framework for robotic reward models that represent task progress as a scalar reward signal. It combines low-rank fine-tuning with token-aligned Plucker rays and synchronous block attention: ray conditioning incorporates camera geometry into visual features and attention queries and keys, while block attention fuses synchronized views inside the pretrained decoder. The framework supports both single-view reward prediction and joint multi-view evaluation through parameter-efficient adaptation of a pretrained Robometer model. On PickCube, single-view adaptation improves progress prediction in every camera group and reduces mean absolute error under a changed field of view by approximately 21% relative to RGB fine-tuning. Across simulated manipulation tasks, joint multi-view prediction reduces progress error by 41-69% compared with averaging single-view RGB predictions and improves temporal ordering in approximately 88% of task-camera groups. On real tasks with fixed and wrist-mounted cameras, mean absolute error decreases by approximately 21% relative to averaged RGB fine-tuning. These results support camera geometry and joint visual evidence as useful components of task-specific robotic reward adaptation.
Comments8 pages, 3 figures, 5 tables