发表机构
Tsinghua University; Beijing Academy of Artificial Intelligence (BAAI); Renmin University of China; Shenzhen Technology University; Hefei University of Technology; Jiangnan University; Chongqing University; The Chinese University of Hong Kong(清华大学; 北京人工智能研究院; 中国人民大学; 深圳技术大学; 合肥工业大学; 江南大学; 重庆大学; 香港中文大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
ChunkTrust将机器人策略的执行视界视为从动作专家证据推断的潜在变量,通过免训练的AHS和可选QHA提升多任务成功率,在RoboTwin2.0和RoboCasa上显著优于基线。
AI 中文摘要
机器人基础策略预测动作块,但在重新规划之前执行多少个动作取决于当前的任务阶段。我们引入了ChunkTrust,它将执行视界视为从动作专家证据中推断出的潜在变量,而非固定的超参数。其免训练的基于动作的视界选择器(AHS)结合了生成轨迹的块内频谱稳定性与已执行历史和预测动作之间的块间连续性。一个带有核遗忘的在线Beta后验在多次重新规划之间跟踪视界偏好。一个轻量级的基于查询的视界适配器(QHA)可选地从互补证据中学习一个上下文条件化的密集先验,并与当前证据和情节本地的Beta记忆融合,而基础策略保持冻结。在RoboTwin2.0和RoboCasa GR1 Tabletop上,AHS提高了每个评估的基础策略配置的整体任务平均成功率,包括在全部50个RoboTwin2.0任务上对$\pi_{0.5}$提高了+6.80个百分点,以及在RoboCasa中对Qwen3GR00T提高了+9.67个百分点。AHS+QHA在八任务$\pi_{0.5}$评估中将相对于Base的增益提高到+9.44个百分点。在四个真实世界家庭任务上,AHS将等任务平均归一化过程得分从50.4%提高到57.5%。消融实验考察了两种证据项、时间记忆和学习到的先验的贡献。项目页面见这个https URL。
英文摘要
Robot foundation policies predict action chunks, but how many actions to execute before replanning depends on the current task phase. We introduce ChunkTrust, which treats the execution horizon as a latent variable inferred from action-expert evidence rather than a fixed hyperparameter. Its training-free Action-aware Horizon Selector (AHS) combines intra-chunk spectral stability of generation traces with inter-chunk continuity between executed history and predicted actions. An online Beta posterior with kernel forgetting tracks horizon preferences across replans. A lightweight Query-based Horizon Adapter (QHA) optionally learns a context-conditioned dense prior from complementary evidence, fused with current evidence and episode-local Beta memory while the base policy remains frozen. Across RoboTwin2.0 and RoboCasa GR1 Tabletop, AHS improves overall task-averaged success for each evaluated base-policy configuration, including gains of +6.80 percentage points on $π_{0.5}$ over all 50 RoboTwin2.0 tasks and +9.67 percentage points on Qwen3GR00T in RoboCasa. AHS+QHA raises the gain over Base to +9.44 percentage points on the eight-task $π_{0.5}$ evaluation. On four real-world household tasks, AHS improves the equal-task mean normalized process score from 50.4% to 57.5%. Ablations examine the contributions of both evidence terms, temporal memory, and the learned prior. Project page is https://hf618.github.io/ChunkTrust.github.io/