arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SurgGMF:用于预期性手术场景渲染的完全因果高斯运动预测

SurgGMF: Fully Causal Gaussian Motion Forecasting for Anticipatory Surgical Scene Rendering

Jingqian Sun, Yichao Tang

arXiv 2609.34733首次发表:更新:

发表机构

Shanghai Research Institute for Intelligent Autonomous Systems, Tongji University; School of Mechanical Engineering, Tongji University; Shanghai Innovation Institute(同济大学上海自主智能无人系统科学中心; 同济大学机械工程学院; 上海创新研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对动态手术场景渲染缺乏未来预测的问题,提出完全因果高斯运动预测框架SurgGMF,通过预测高斯运动状态实现预期性渲染,在12个视频片段上优于经典动力学基线。

AI 中文摘要

动态手术场景建模对于机器人感知、仿真和决策支持至关重要。尽管现有的神经渲染方法能够高效地重建和渲染可变形手术场景,但它们主要侧重于观测帧的重建,而非预测未来场景状态。为此,我们提出了SurgGMF,一个用于预期性手术场景渲染的完全因果高斯运动预测框架。SurgGMF不直接预测未来的RGB图像,而是从历史高斯运动场中预测由位置、尺度和旋转残差(X/S/R)表示的未来高斯运动状态。为防止目标泄漏,我们引入了一种全因果最后渲染协议,在该协议下,未来高斯状态在渲染时不访问目标帧的高斯属性,同时保持因果外观传播。我们在12个EndoNeRF和StereoMIS视频片段上,使用神经时序学习器和经典动力学基线,在统一预测协议下评估了SurgGMF。在渲染空间中,学习到的高斯运动预测始终优于经典动力学基线,展现出超越手工状态外推的性能提升。延迟分析进一步揭示了精度与效率之间的权衡:在当前实现下,TKAN实现了最高精度,而GRU和LSTM提供了更有利的模块级延迟特性。这些结果使SurgGMF成为可复现的因果高斯运动预测框架,并将手术高斯表示从回顾性重建推进到预测性场景建模。

英文摘要

Dynamic surgical scene modeling is essential for robotic perception, simulation, and decision support. Although existing neural rendering methods enable efficient reconstruction and rendering of deformable surgical scenes, they remain primarily focused on observed-frame reconstruction rather than forecasting future scene states. To this end, we present SurgGMF, a fully causal Gaussian motion forecasting framework for anticipatory surgical scene rendering. Rather than predicting future RGB images directly, SurgGMF forecasts future Gaussian motion states represented by position, scale, and rotation residuals (X/S/R) from historical Gaussian motion fields. To prevent target leakage, we introduce a full-causal-last rendering protocol, where future Gaussian states are rendered without accessing target-frame Gaussian attributes while preserving causal appearance propagation. We evaluate SurgGMF on 12 EndoNeRF and StereoMIS video slices using neural temporal learners and classical dynamics baselines under a unified forecasting protocol. Learned Gaussian motion forecasting consistently outperforms classical dynamics baselines in render space, demonstrating gains beyond hand-crafted state extrapolation. Latency analysis further reveals an accuracy--efficiency trade-off: under the current implementations, TKAN achieves the highest accuracy, whereas GRU and LSTM provide more favorable module-level latency profiles. These results establish SurgGMF as a reproducible framework for causal Gaussian motion forecasting and advance surgical Gaussian representations from retrospective reconstruction toward predictive scene modeling.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑