arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.33075cs.AI

QureRadEmbed:通过属性与推理监督构建放射学相似性结构

QureRadEmbed: Structuring Radiological Similarity through Attribute and Reasoning Supervision

Janhavi Prabhu, Sahil, Shivam Ashok Shukla, Manoj Tadepalli

首次发表
浏览论文内容

中文总结 AI 辅助

QureRadEmbed通过属性与推理监督训练4B放射学编码器,显著提升相似性排序和检索性能,支持高效报告比较与结构化索引。

中文摘要 AI 辅助

放射学相似性取决于疾病关系以及诸如侧别、肺叶、严重程度、大小和确定性等细节。广泛的生物医学相似性可能忽视这些限定词,尤其是当多个属性同时变化时。我们引入了QureRadEmbed,一个4B规模的放射学感知编码器,通过两种互补信号进行训练:RadSim提供确定性的、属性分解的排序目标,而RadThought将报告与分层证据和推理描述对齐。一个三阶段课程将这些信号与报告三元组、扰动发现以及单属性和跨属性对比相结合。最终模型在十个受控合成属性上达到了0.996的平均排序准确率,并将与设计联合属性目标的Spearman相关性从0.501提升至0.976。在外部发现到印象检索中,Recall@1在Open-I上达到10.5%,在测试XR上达到12.4%,在测试CT上达到42.9%,而基线的相应数值分别为6.6%、5.7%和31.4%。冻结嵌入仅使用100份带标签的测试XR报告即可支持发现提取(宏F1为0.481,而基线为0.412)。在我们的基准测试中,整份报告比较每1000对耗时8.8秒,而生成式评估器GREEN则需2755.1秒。句子级比较提高了对局部差异的敏感性,尽管在若干专家评分和细微错误任务上,生成式评估仍然更强。这些结果支持可复用的放射学感知表示,用于搜索、结构化报告索引和高效报告比较。

英文摘要

Radiological similarity depends on disease relationships and on fine details such as laterality, lobe, severity, size, and certainty. Broad biomedical similarity can overlook these qualifiers, particularly when several attributes vary together. We introduce QureRadEmbed, a 4B radiology-aware encoder trained with two complementary signals: RadSim supplies deterministic, attribute-decomposed ranking targets, while RadThought aligns reports with hierarchical evidence and reasoning descriptions. A three-stage curriculum combines these signals with report triplets, finding perturbations, and single- and cross-attribute contrasts. The final model achieves 0.996 mean ordering accuracy across ten controlled synthetic attributes and raises Spearman correlation with the designed joint-attribute targets from 0.501 to 0.976. On external findings-to-impression retrieval, Recall@1 reaches 10.5% on Open-I, 12.4% on testing XR, and 42.9% on testing CT, compared with 6.6%, 5.7%, and 31.4% for its backbone. Frozen embeddings support finding extraction with only 100 labeled testing-XR reports (macro-F1 0.481 versus 0.412 for the backbone). Whole-report comparison costs 8.8 seconds per 1,000 pairs in our benchmark, versus 2,755.1 seconds for the generative evaluator GREEN. Sentence-level comparison improves sensitivity to local discrepancies, although generative evaluation remains stronger on several expert-rated and subtle-error tasks. The results support reusable radiology-aware representations for search, structured report indexing, and efficient report comparison.

补充信息

↑