arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

templar:从大规模临床语料中智能体式归纳与演化标准化放射学报告模板

templar: agentic induction and evolution of standardized radiology reporting templates from large-scale clinical corpora

Xiaotian Hu, Mingxuan Liu, Zhonghan Wang, Xinfeng Zhang, Yiming Huang, Ziang Wang, Kasidit Anmahaepong, Yijin Li, Yifei Chen, Hongjia Yang, Zihan Li, Qiyuan Tian

arXiv 2610.05247首次发表:更新:

发表机构

Tsinghua University; University of Hong Kong; University of California, San Diego(清华大学; 香港大学; 加利福尼亚大学圣迭戈分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对现有放射学报告模板归纳方法受上下文限制、缺乏外部依据和适应性不足的问题,提出TEMPLAR智能体框架,通过三个智能体协作从大规模临床语料中归纳并演化标准化模板,在四个数据集上优于现有方法,提升覆盖率、信息保真度和诊断保真度。

AI 中文摘要

结构化放射学报告减轻了自由文本报告的异质性,但其益处取决于高质量的报告模板。在实践中,此类模板通常通过劳动密集型的专家共识构建,因此在不同机构间存在差异,且滞后于不断发展的临床实践。大型语言模型(LLMs)实现了模板的自动归纳,但现有方法仍然有限:单一LLM归纳受限于上下文长度,而语料库规模方法ASTAR产生的是静态的、封闭语料库的模板,缺乏外部依据或下游适应性。为解决这些局限,我们提出TEMPLAR,一种以模板为中心的智能体框架,用于从大规模临床语料中归纳和演化标准化放射学报告模板。TEMPLAR将模板视为持久的核心状态,与两个具有溯源意识的知识图谱并存,即约束模板构建的解剖图谱和支持从发现到诊断推理的诊断图谱。三个智能体在此状态上运作。归纳智能体通过双视图相似性聚类,从受解剖约束的Span-Triple原子中推导出规范的临床槽位;演化智能体随后将这些槽位组装成分层模板,并在一致性约束、外部临床证据和下游结构化反馈下对其进行修订;临床智能体将演化后的模板应用于报告结构化、重建和诊断推理。在四个数据集上,TEMPLAR在覆盖率、信息保真度和诊断保真度方面优于ASTAR、三个医学LLM和六个通用LLM,同时获得最高或并列最高的LLM评定的模板质量。其相对于ASTAR的保真度优势在跨数据集迁移中持续存在,累积消融实验支持其关键组件的互补贡献。

英文摘要

Structured radiology reporting mitigates the heterogeneity of free-text reports, yet its benefits depend on high-quality reporting templates. In practice, such templates are conventionally built through labor-intensive expert consensus and therefore vary across institutions and lag behind evolving clinical practice. Large language models (LLMs) enable automated template induction, but existing approaches remain limited: single-LLM induction is constrained by context length, and the corpus-scale method ASTAR produces a static, closed-corpus template without external grounding or downstream adaptation. To address these limitations, we propose TEMPLAR, a TEMPLate-centric Agentic framework for inducing and evolving standardized Radiology reporting templates from large-scale clinical corpora. TEMPLAR treats the template as a persistent central state maintained alongside two provenance-aware knowledge graphs, namely an anatomical graph that constrains template construction and a diagnostic graph that supports finding-to-diagnosis reasoning. Three agents operate on this state. The Induction Agent derives canonical clinical slots from anatomy-constrained Span-Triple atoms via dual-view similarity clustering; the Evolution Agent then assembles these slots into a hierarchical template and revises it under consistency constraints, external clinical evidence, and downstream structuring feedback; and the Clinical Agent applies the evolved template to report structuring, reconstruction, and diagnostic reasoning. Across four datasets, TEMPLAR outperforms ASTAR, three medical LLMs, and six general-purpose LLMs in coverage, information fidelity, and diagnostic fidelity, while achieving the highest or tied-highest LLM-rated template quality. Its fidelity advantages over ASTAR persist under cross-dataset transfer, and cumulative ablations support complementary contributions of its key components.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑