arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.15009cs.ROcs.CV

ForceU-VLA:用于实体超声扫描的力感知视觉-语言-动作模型

ForceU-VLA: A Force-Aware Vision-Language-Action Model for Embodied Ultrasound Scanning

Xingzheng Wu, Cheng Zhang, Guihao Yan, Xifeng Hu, Zhi Liu, Qing Cai

首次发表
浏览论文内容

中文总结 AI 辅助

本文针对现有实体超声扫描方法的不足,提出力感知视觉-语言-动作模型ForceU-VLA,设计FUSFM与SAMM模块,构建ForceU-VLA-Data数据集,实验证实其可提升超声扫描的接触稳定性与压力调节能力。

中文摘要 AI 辅助

实体智能超声扫描通过整合感知、决策与执行能力,实现超声检查流程的自动化与标准化。然而,现有方法存在力与超声模态间建模松散、缺乏扫描阶段感知的问题,限制了其捕捉探头-组织动态交互的能力。为解决这些问题,本文提出ForceU-VLA,一种用于自主实体超声扫描的力感知视觉-语言-动作模型,其在整个扫描过程中利用力信号与超声图像反馈,实现准确、高质量的超声采集。首先,本文提出力-超声协同融合模块(FUSFM),协同融合超声视觉与力反馈信息,为探头运动提供稳定、可靠的引导。其次,提出阶段自适应调制机制(SAMM),通过自适应调制多模态特征以提升其表征质量,适配不同扫描阶段的任务需求。此外,本文引入ForceU-VLA-Data,一个真实世界的力感知实体超声数据集,整合视觉、力与动作信号,包含两个器官的五种代表性临床扫描视图数据,由专家采集的450条轨迹,约10万帧同步多模态帧。大量实验结果表明,ForceU-VLA显著提升了实体超声扫描中的接触稳定性与探头压力调节能力,从而有效提高任务执行质量与整体系统可靠性。源代码可在指定URL获取。

英文摘要

Embodied intelligent ultrasound scanning enables the automation and standardization of the ultrasound examination process by integrating perception, decision-making, and execution capabilities. However, existing methods suffer from loosely coupled modeling between force and ultrasound modalities and lack awareness of scanning stages, which limits their ability to capture dynamic probe-tissue interactions. To address these issues, we propose ForceU-VLA, a force-aware Vision-Language-Action model for autonomous embodied ultrasound scanning, which leverages force signals and ultrasound image feedback throughout the scanning process to enable accurate and high-quality ultrasound acquisition. Firstly, we propose a Force-Ultrasound Synergistic Fusion Module (FUSFM) that synergistically fuses ultrasound visual and force-feedback information to provide stable, reliable guidance for probe motion. Secondly, a Stage-Adaptive Modulation Mechanism (SAMM) is proposed to accommodate the task requirements across different scanning stages by adaptively modulating multimodal features to enhance their representation quality. Additionally, we introduce ForceU-VLA-Data, a real-world, force-aware embodied ultrasound dataset that integrates visual, force, and action signals, including data from two organs across five representative clinical scanning views, and comprising 450 expert-collected trajectories with approximately 100,000 synchronized multimodal frames. Extensive experimental results demonstrate that ForceU-VLA significantly improves contact stability and probe pressure regulation in embodied ultrasound scanning, thereby effectively enhancing task execution quality and overall system reliability. The source code is available at https://github.com/VMVLab/ForceU-VLA.

发表机构

  • Faculty of Computer Science and Technology, Ocean University of China(中国海洋大学计算机科学与技术学院)
  • School of Information Science and Engineering, Shandong University(山东大学信息科学与工程学院)
  • Innovation School of Artificial Intelligence, Hefei University of Technology(合肥工业大学人工智能创新学院)

机构由 AI 辅助整理,请以论文原文为准。

↑