发表机构
University of Macau; University of Macau Advanced Research Institute in Hengqin; The Chinese University of Hong Kong; The Hong Kong Polytechnic University; Shanghai Jiao Tong University; Tongren Hospital; Duke University(澳门大学; 澳门大学横琴研究院; 香港中文大学; 香港理工大学; 上海交通大学; 同仁医院; 杜克大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出初步联合视觉-轨迹世界动作模型,从历史手术观测中同时预测未来视觉状态与器械轨迹,采用分块自回归展开法,提升了预测性能,证明了联合预测的可行性,但仍存在长范围预测的退化与误差挑战。
AI 中文摘要
可靠的外科手术规划要求模型不仅要预测器械如何运动,还要预测手术视觉状态如何随这种运动演变。现有方法通常将未来场景生成和器械轨迹预测视为两个独立任务:仅场景模型无法在轨迹层面直接评估未来器械运动的准确性,而仅轨迹模型无法捕捉器械运动的视觉后果,导致预测轨迹与未来场景演变之间的一致性问题未得到解决。联合预测两者可通过实现明确的轨迹层面评估,同时建模相应的视觉演变,更完整地描述手术动作-场景动态。为弥合这一差距,本文提出一种初步的联合视觉-轨迹世界动作模型,可从历史手术观测中同时预测未来视觉状态和器械轨迹。具体而言,我们将历史视频帧和工具轨迹编码为潜在表示,经时空编码器处理后,通过独立的视觉状态和轨迹预测头进行解码。基于此初步架构,采用分块自回归展开法重复预测15个未来步骤。在所有评估的预测范围内,分块策略的性能始终优于直接一次性预测,将第一片段的峰值信噪比(PSNR)从18.86 dB提升至23.11 dB,将平均位移误差(ADE)从45.77像素降低至22.22像素。这些结果证明了联合视觉-运动预测的初步可行性,但我们观察到在较长预测范围内存在渐进式视觉退化和累积轨迹误差,这些仍是外科世界动作建模未来需解决的重要挑战。
英文摘要
Reliable surgical planning requires models to anticipate not only how instruments will move, but also how the operative visual state will evolve together with such motion. Existing approaches typically treat future scene generation and instrument trajectory prediction as two separate tasks. Scene-only models cannot directly evaluate the accuracy of future instrument motion at the trajectory level, while trajectory-only models fail to capture the visual consequences of instrument movement, leaving the consistency between predicted trajectories and future scene evolution unaddressed. Jointly forecasting both provides a more complete account of surgical action-scene dynamics by enabling explicit trajectory-level evaluation while simultaneously modeling the corresponding visual evolution. To bridge this gap, we present a preliminary joint visual-trajectory world-action model that simultaneously forecasts future visual states and instrument trajectories from historical surgical observations. Specifically, we encode historical video frames and tool trajectories into latent representations, which are processed by a temporal-spatial encoder and subsequently decoded through separate visual-state and trajectory prediction heads. Based on this preliminary architecture, a chunked autoregressive rollout is repeatedly applied to predict fifteen future steps. The chunked strategy consistently outperforms direct one-shot prediction across all evaluated horizons, improving first-segment PSNR from 18.86 to 23.11 dB and reducing ADE from 45.77 to 22.22 pixels. These results demonstrate the initial feasibility of joint visual-motion forecasting. However, we observe progressive visual degradation and accumulated trajectory errors over longer prediction horizons, which remain important challenges for future surgical world-action modeling.