发表机构
University of Missouri; Missouri Cancer Registry and Research Center; MU Institute for Data Science and Informatics(密苏里大学; 密苏里癌症登记与研究中心; 密苏里大学数据科学与信息学研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
CRISS是一种检索增强生成聊天机器人,利用领域知识库和LLM为癌症登记员提供有引文支持的回答,实验表明其优于非RAG方法,并保留人工监督。
AI 中文摘要
癌症登记员,包括肿瘤数据专家(ODS),必须解读复杂且频繁更新的编码和分期标准。我们开发了CRISS(癌症登记智能支持系统),这是一种检索增强生成(RAG)对话助手,可提供快速、有引文支持的登记指南访问。本研究评估了CRISS是否能够(1)支持准确且有引文支持的响应,(2)改善对相关指南的访问和解读,以及(3)支持培训/帮助台使用,同时保留对最终抽象决策的人工监督。我们从国家癌症登记标准构建了一个领域特定的知识库,将其分割为带元数据标签的段落,并索引为密集嵌入。检索到的段落通过大型语言模型(LLM)生成有引文依据的响应。对Gemini和GPT系列中的开放权重、专有和非RAG基线模型,使用LLM作为评判者的协议,在简单、中等和困难的登记问题上进行了评估。RAG配置始终优于非RAG方法,尤其是在问题难度增加时。RAG的平均基础性得分在简单/中等/困难层级分别为0.62/0.56/0.59,而非RAG为0.29/0.26/0.29。RAG模型总体上还获得了更高的语义相似度得分。专有RAG模型在简单和中等问题上表现最强,而本地RAG模型在困难问题上排名最高,专有模型通常更为谨慎。领域特定的RAG改善了癌症登记问题的证据基础和响应质量,同时在不同复杂度级别上提供了有引文支持的辅助。CRISS展示了以人为中心、有引文支持的AI在支持癌症登记员方面的潜力,同时保留了对最终编码决策的人工监督。
英文摘要
Cancer registrars, including Oncology Data Specialists (ODSs), must interpret complex and frequently updated coding and staging standards. We developed CRISS (Cancer Registry Intelligent Support System), a retrieval-augmented generation (RAG) conversational assistant that provides rapid, citation-supported access to registry guidance. This study evaluated whether CRISS could (1) support accurate and citation-supported responses, (2) improve access to and interpretation of relevant guidance, and (3) support training/helpdesk use while preserving human oversight of final abstraction decisions. We built a domain-specific knowledge base from national cancer registry standards, segmented into metadata-tagged passages and indexed as dense embeddings. Retrieved passages were used to generate citation-grounded responses through a large language model (LLM). Open-weight, proprietary, and non-RAG baseline models across Gemini and GPT families were evaluated on easy, medium, and hard registry questions using an LLM-as-a-Judge protocols. RAG configurations consistently outperformed non-RAG approaches, especially as question difficulty increased. Mean grounding scores for RAG were 0.62/0.56/0.59 across easy/medium/hard tiers versus 0.29/0.26/0.29 for non-RAG. RAG models also achieved higher semantic-similarity scores overall. Proprietary RAG models performed strongest on easy and medium questions, while local RAG models ranked highest on hard questions and proprietary models were generally more cautious. Domain-specific RAG improved evidence grounding and response quality for cancer registry questions while enabling citation-supported assistance across complexity levels. CRISS demonstrates the potential of human-centered, citation-grounded AI to support cancer registrars while preserving human oversight for final coding decisions.
Comments21 pages, 13 figures, 7 tables. Keywords: cancer registry, retrieval-augmented generation, large language models, conversational AI, clinical informatics, oncology data specialists, medical question answering, AI safety, clinical decision support