arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.13238cs.CLcs.CV

锥形束CT报告生成中部分可观测目标下的临床推理

Clinical Reasoning Under a Partially Observed Objective in Cone Beam CT Report Generation

  • DiceMed
  • University of Maryland(马里兰大学)
  • Indira Gandhi National Open University(英迪拉·甘地国立开放大学)
  • Indian Institute of Technology Jammu(印度理工学院贾姆穆分校)

机构由 AI 辅助整理,请以论文原文为准。

Ajo Babu George, Govind Arun, Sidharth N Krishna, Uma Ranjan

中文总结 AI 辅助

本研究针对锥形束CT报告生成,提出部分可观测复合目标优化方法,发现词汇指标奖励听写惯例而非解剖结构,并交付了结合几何约束的报告系统。

中文摘要 AI 辅助

本文对锥形束计算机断层扫描生成颌面部报告的任务进行了评分,评分采用复合目标,其中80%的权重分配给基于大语言模型的事实蕴含判断,20%分配给词汇重叠,而在开发过程中仅词汇重叠这一部分(即五分之一)是可见的。评分器的BLEU-4和METEOR例程以纯Python复现,并与参考实现达到机器精度一致;一个离线蕴含替代模型(其区分针对一位患者撰写的报告与针对另一位患者撰写的报告的能力,曲线下面积达0.987)使得复合目标的直接优化变得足够廉价。在包含622例病例的公开数据集中,根据可见的词汇排名选择的报告得分为0.2909,而根据复合目标选择的报告得分为0.4122,这是因为追求n-gram重叠导致蕴含精确率从0.522降至0.266。一个在该数据集上微调的2900万参数编码器,在985条陈述上的患病率加权折外曲线下面积为0.486,与语料库先验无法区分;而从图像头读取的九个数字在下颌骨覆盖方面达到0.945,在髁突覆盖方面达到0.872;仅采集中心预测句子选择的能力为0.718,而基于图像的模型为0.663,这表明词汇指标所奖励的是听写惯例而非解剖结构。所交付的系统输出八条无条件陈述和五条受限于极性、侧别和牙齿级别一致性约束的基于图像头几何的陈述,并在来自未见中心的50例保留病例上达到METEOR 0.3542。数据集和代码可在以下网址获取:此https URL

英文摘要

Maxillofacial report generation from cone beam computed tomography is scored here by a composite objective placing 80% of its weight on a large language model judgement of factual entailment and 20% on lexical overlap, of which only the lexical fifth is visible during development. The grader's BLEU-4 and METEOR routines are reproduced in pure Python and match the reference to machine precision, and an offline entailment surrogate, which tells a report written for one patient from one written for another at an area under the curve of 0.987, makes the composite objective cheap enough to optimise directly. Over the 622-case public release, a report selected against the visible lexical ranking scores 0.2909, whereas one selected against the composite objective scores 0.4122, because pursuing n-gram overlap drives entailment precision from 0.522 down to 0.266. A 29 million parameter encoder fine-tuned on the release reaches a prevalence-weighted out-of-fold area under the curve of 0.486 over 985 statements, indistinguishable from the corpus prior, while nine numbers read from the image header reach 0.945 for mandible coverage and 0.872 for condyle coverage, and acquisition centre alone predicts sentence choice at 0.718 against 0.663 for the image-derived model, identifying dictation convention rather than anatomy as the quantity the lexical metrics reward. The delivered system emits eight unconditional statements and five gated on header geometry under polarity, laterality and tooth-level consistency constraints, and reaches METEOR 0.3542 over 50 held-out cases from an unseen centre. The dataset and code are available at https://github.com/GIND123/CBCT-Clinical-Reasoner

补充信息

↑