arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.11610cs.CVcs.AI

超越报告模仿:基于可见证据的临床感知多图像超声报告生成

Beyond Report Imitation: Clinically Aware Multi-Image Ultrasound Report Generation from Visible Evidence

Yuchen Yang, Xin Wang, Lufan Wang, Yinghong Pan, Yujuan Feng, Yuqing Yang

首次发表
浏览论文内容

中文总结 AI 辅助

针对多图像超声报告生成的临床行为对齐问题,提出CAMEO框架,构建相关数据集并在基准上显著提升指标,凸显证据基础监督等的价值。

中文摘要 AI 辅助

从多幅图像生成超声报告需要聚合不同视图的临床证据,但存档的关键帧仅捕获了动态检查的部分内容。原始报告模仿因此与视觉监督不一致:对于完整检查具有临床有效性的内容,可能无法从模型可用的图像中验证,这一差距造成了临床行为对齐问题。模型必须保留可见发现、避免诊断反转和无依据的补全,且不能坍缩为保守模板。我们提出CAMEO——一种用于超声报告生成的临床感知多图像证据基础编排框架。阶段I学习超声视觉语言原语;阶段II通过将可信的可见报告点提炼为多图像问答和报告风格监督来执行跨视图证据基础;阶段III使用面向临床错误的偏好对执行临床感知偏好对齐。我们从USReport构建了包含17670个证据基础配对图像训练实例的USReport-Distilled,以及包含21869个偏好对的USReport-Pref;还使用25631个PubMedVision-US超声指令样本进行领域适应和多图像指令调优。在主要的USReport-Distilled基准上,CAMEO将BLEU-1从0.25提升至0.40,ROUGE-1从0.28提升至0.45,METEOR从0.27提升至0.43,同时将ClinicalScore从55.02提升至74.20。这些结果凸显了证据基础监督、临床感知对齐以及临床结构化评估对于可靠超声报告生成的价值。

英文摘要

Generating ultrasound reports from multiple images requires aggregating clinical evidence across views, yet archived key frames capture only part of the dynamic examination. Raw-report imitation is therefore misaligned with visual supervision: content that is clinically valid for the full examination may be unverifiable from the images available to a model. This gap creates a clinical behavior alignment problem. A model must preserve visible findings, avoid diagnostic reversals and unsupported completion, and not collapse into conservative templates. We propose CAMEO, a Clinically Aware Multi-image Evidence-grounded Orchestration framework for ultrasound report generation. Stage I learns ultrasound visual-language primitives; Stage II performs Cross-View Evidence Grounding by distilling trusted visible report points into multi-image QA and report-style supervision; and Stage III performs Clinically Aware Preference Alignment using clinical-error-oriented preference pairs. From USReport, we construct USReport-Distilled with 17,670 evidence-grounded paired-image training instances and USReport-Pref with 21,869 preference pairs; we additionally use 25,631 PubMedVision-US ultrasound instruction samples for domain adaptation and multi-image instruction tuning. On the primary USReport-Distilled benchmark, CAMEO improves over EchoVLM from 0.25 to 0.40 BLEU-1, 0.28 to 0.45 ROUGE-1, and 0.27 to 0.43 METEOR, while raising ClinicalScore from 55.02 to 74.20. These results underscore the value of evidence-grounded supervision, clinically aware alignment, and clinically structured evaluation for reliable ultrasound report generation.

发表机构

  • State Key Laboratory of Networking and Switching Technology, Beijing University of Posts and Telecommunications(网络与交换技术国家重点实验室,北京邮电大学)
  • College of Computer Science, Beijing University of Technology(北京工业大学 计算机学院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑