arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

摊销反馈规划:将基于模型的展开转化为可执行策略

Amortized Feedback Planning: Turning Model-Based Rollouts into Executable Policies

Jeonggyu Huh

arXiv 2609.35012首次发表:更新:

发表机构

SKKU(韩国科学技术院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出摊销反馈规划方法,通过贝尔曼梯度细化将模型展开转化为可执行策略,在非LQG和推力任务中显著降低计算成本并保持性能。

AI 中文摘要

闭环规划考虑了未来依赖于观测的动作,但在部署时重复执行可能代价高昂。贝尔曼梯度(BG)细化通过现有actor区分条件展开,修正当前动作,并将结果存储在可执行策略中。反向扫描重用已部署的未来反馈;在已识别的高斯信念模型中的全水平展开无需学习价值评论器。非线性误差递归将二次连续误差、局部修正和策略存储联系起来。一个受控的非LQG示例展示了具有一致存储的二阶策略精度,而理想仿射LQG允许精确的反向恢复。在非二次推力中,BG将actor成本降低了2.47-3.20%,并保持在测试反馈MPC的0.13-0.53%以内;原始策略在seed-0原生审计中以5-6微秒执行。在匹配存储下,BG在单GPU上达到竞争性工厂成本,其构建时间约为反馈教师加学生构建时间的九分之一(机械臂)和六十分之一(对接),不包括共享学习和节点准备。更丰富的公共地图显著缩小了一些学生-BG差距。这些结果揭示了学习反馈、局部修正和存储在可执行控制中的作用。

英文摘要

Closed-loop planning accounts for future observation-dependent actions but can be expensive to repeat at deployment. Bellman-gradient (BG) refinement differentiates conditional rollouts through an existing actor, corrects the current action, and stores the result in an executable policy. A backward sweep reuses deployed future feedback; full-horizon rollouts in an identified Gaussian belief model need no learned value critic. A nonlinear error recursion links quadratic continuation error, local correction, and policy storage. A controlled non-LQG example exhibits second-order policy accuracy with consistent storage, while ideal affine LQG admits exact backward recovery. In nonquadratic thrust, BG reduces actor cost by 2.47-3.20% and remains within 0.13-0.53% of the tested feedback MPC; original policies execute in 5-6 microseconds in a seed-0 native audit. With matched storage, BG attains competitive plant costs at about one ninth (arm) and one sixtieth (docking) of feedback-teacher-plus-student construction time on one GPU, excluding shared learning and node preparation. Richer common maps substantially narrow some student-BG gaps. These results expose the roles of learned feedback, local correction, and storage in executable control.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑