arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.13103cs.ROcs.SYeess.SY

S2-HWM:用于长程外科机器人操作的稀疏事件结构化分层世界模型

S2-HWM: Sparse Event-Structured Hierarchical World Model for Long-Horizon Surgical Robot Manipulation

Shuzhe Zhang, Xin Zhu, Yinling Qian, Qiong Wang

首次发表
浏览论文内容

中文总结 AI 辅助

针对长程外科机器人操作奖励稀疏、进展隐含的问题,提出S2-HWM模型,通过事件证据协调管理器与工作器,在SurRoL的PegTransfer任务上成功率达98.7%,优于基准模型。

中文摘要 AI 辅助

长程外科机器人操作极具挑战性,因为任务奖励十分稀疏,而有意义的交互变化却以不规则的间隔发生。现有的世界模型智能体通常以原始步长分辨率进行想象,使得持续时间可变的任务进展隐含其中。手动指定的阶段可以提供中间结构,但其特定于任务的边界难以与依赖状态的交互转换对齐。我们提出S2-HWM,即一种稀疏事件结构化分层世界模型,它从原始潜轨迹中学习稀疏事件证据,以协调事件级管理器和原始步长工作器。事件证据会调度管理器的目标更新,且每个选定的潜目标会条件化工作器的原始动作,直至下一次更新。学习到的事件证据还会为事件转换模型(ETM)形成持续时间可变的片段,ETM会预测下一个边界随机状态、片段持续时间以及累积的片段奖励。将这些事件级预测连接起来,可为管理器学习提供超出原始想象 horizon 的持续时间可变的延续,而工作器则保留原始步长的演员-评论家学习。在基于SurRoL的PegTransfer任务上,S2-HWM达到了98.7%的成功率,比平坦的GAS DreamerV3基准高出22.7个百分点。

英文摘要

Long-horizon surgical robot manipulation is challenging because task rewards are sparse, while meaningful interaction changes occur at irregular intervals. Existing world-model agents typically imagine at primitive-step resolution, leaving variable-duration task progress implicit. Manually specified stages can provide intermediate structure, but their task specific boundaries are difficult to align with state-dependent interaction transitions. We propose S2-HWM, a Sparse Event-Structured Hierarchical World Model that learns sparse event evidence from primitive latent trajectories to coordinate an event-level manager and a primitive-step worker. The event evidence schedules manager goal updates, and each selected latent goal conditions the worker's primitive actions until the next update. The learned event evidence also forms variable-duration segments for an Event Transition Model (ETM), which predicts the next?boundary stochastic state, segment duration, and accumulated segment reward. Chaining these event-level predictions provides a variable-duration continuation beyond the primitive imagination horizon for manager learning, while the worker retains primitive-step actor-critic learning. On a SurRoL-based PegTransfer task, S2-HWM achieves a success rate of 98.7%, outperforming the flat GAS DreamerV3 baseline by 22.7 percentage points.

↑