arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.12590cs.AIcs.CV

用于循证甲状腺超声诊断与报告的可审计智能体AI

Auditable agentic AI for evidence-grounded thyroid ultrasound diagnosis and reporting

  • School of Computer Science and Engineering, Sun Yat-sen University(中山大学计算机科学与工程学院)
  • School of Software Engineering, Sun Yat-sen University(中山大学软件工程学院)
  • School of Mathematical Sciences, Zhejiang University(浙江大学数学科学学院)
  • Perelman School of Medicine, University of Pennsylvania(宾夕法尼亚大学佩雷尔曼医学院)
  • Zhujiang Hospital, Southern Medical University(南方医科大学珠江医院)
  • College of Mathematical Medicine, Zhejiang Normal University(浙江师范大学数学医学院)

机构由 AI 辅助整理,请以论文原文为准。

Haifan Gong, Shiyu Chen, Bodong Wang, Yuqi Wang, Shijie Wang, Guoliang You, Xinyu Xiong, Haowei Wang, Mingzhi Mao, Dexing Kong, Qinghua Liu, Wei Lou, Fei Chen, Guanbin Li

中文总结 AI 辅助

本文提出可审计的智能体AI系统ThyroidXAgent,基于多中心数据集开发,可协同完成甲状腺超声诊断多任务,提升分类准确率与报告一致性,缩短耗时,支持临床医生修正。

中文摘要 AI 辅助

甲状腺超声诊断需要协同完成病灶定位、测量、风险分层及报告生成,然而多数AI系统孤立处理这些任务,对临床审阅的支持有限。本文提出ThyroidXAgent,一种可与临床医生交互的智能体AI系统,它能协调专用诊断工具并将其输出存储为可审计的病例级证据记录。该系统基于OpenThyroidDB开发,这是一项多中心多任务资源,整合了约30万张超声图像及2.4万份配对报告,在28458个非重叠测试病例上进行评估,其中包括来自NHC-MISD-TUS私人队列35个中心的8721个病例。在异质性数据集上,ThyroidXAgent的结节分割平均Dice评分为87.21%,良恶性分类平均AUROC为0.9466;同一工作流还支持淋巴结转移预测及滤泡性与乳头状甲状腺癌分类,AUROC分别为0.864和0.805。在报告生成方面,循证组装在三个队列上的表现优于多模态语言模型基线。本文引入的病灶级临床语义指标ThyClinScore与位置感知语言模型评判者的相关性最强。ThyroidXAgent提升了医生分类准确率,使报告诊断一致性从70.3%升至86.2%,并将分割与报告时间分别缩短35.9%和27.4%。这些结果支持开发可审计、可由临床医生修正的智能体AI用于甲状腺超声诊断与报告。

英文摘要

Thyroid ultrasound diagnosis requires coordinated lesion localization, measurement, risk stratification and reporting, yet most AI systems address these tasks in isolation and provide limited support for clinical review. We present ThyroidXAgent, a clinician-interactive agentic AI system that coordinates specialized diagnostic tools and stores their outputs as an auditable case-level evidence record. The system was developed using OpenThyroidDB, a multicentre, multitask resource integrating approximately 0.3 million ultrasound images and 24,000 paired reports, and was evaluated on 28,458 non-overlapping test cases, including 8,721 cases from 35 centres in the private NHC-MISD-TUS cohort. Across heterogeneous datasets, ThyroidXAgent achieved a mean Dice score of 87.21 percent for nodule segmentation and a mean AUROC of 0.9466 for benign-malignant classification. The same workflow supported lymph-node metastasis prediction and follicular versus papillary thyroid carcinoma classification, with AUROCs of 0.864 and 0.805, respectively. For report generation, evidence-grounded assembly outperformed multimodal language-model baselines across three cohorts. ThyClinScore, a lesion-level clinical semantic metric introduced here, showed the strongest correlation with a location-aware language-model judge. ThyroidXAgent improved physician classification accuracy, increased report diagnostic consistency from 70.3 percent to 86.2 percent, and reduced segmentation and reporting time by 35.9 percent and 27.4 percent, respectively. These findings support auditable, clinician-correctable agentic AI for thyroid ultrasound diagnosis and reporting.

补充信息

↑