跨语言临床标注投影作为约束文本生成:一项六语言研究
Cross-Lingual Clinical Annotation Projection as Constrained Text Generation: A Six-Language Study
浏览论文内容
中文总结 AI 辅助
本研究提出将跨语言临床标注投影作为约束文本生成任务,开发直接LLM投影工作流,在六语言评估中显著提升严格F1分数,有效减少专家标注成本。
中文摘要 AI 辅助
背景:旨在确定跨语言临床标注投影是否可以表述为一种保留文本的、文档级生成任务,该任务能为多语言临床语料库构建生成可验证的字符级标注,并刻画其相对于基于候选的投影流程的稳健性和计算权衡。方法:我们开发了一种约束LLM投影工作流,将实体标签直接插入不可变的目标语言文本中,随后进行确定性验证和字符偏移重建。我们将其与监督候选跨度投影和混合机器学习-大语言模型精炼方法进行了评估,用于将西班牙语疾病、症状和程序标注迁移到六种语言。评估使用MultiClinAI金标准,采用严格跨度匹配和字符重叠F1分数。结果:直接LLM投影取得了最强且最一致的性能。GLM 5.2在18个语言-实体组合上的平均严格F1为0.9201,而本地可部署的Gemma4:31B达到了0.9133。最佳LLM配置在所有18个设置中将严格F1相对于先前最先进水平提高了0.0564至0.1512,产生了55,416个具有重建偏移的基于语料库的提及。结论:基于直接LLM的投影能够实现高质量的多语言临床标注迁移,并为将临床NLP资源扩展到标注数据集和语言特定工具较少的语言提供了一种实用方法。结合本地推理和确定性验证,它可以大幅减少多语言临床语料库构建的专家时间和成本。
英文摘要
Background: To determine whether cross-lingual clinical annotation projection can be formulated as a text-preserving, document-level generative task that produces verifiable character-level annotations for multilingual clinical corpus construction, and to characterize its robustness and computational trade-offs relative to candidate-based projection pipelines. Methods: We developed a constrained LLM projection workflow that inserts entity tags directly into immutable target-language text, followed by deterministic validation and character-offset reconstruction. We evaluated it alongside supervised candidate-span projection and hybrid ML-LLM refinement for transferring Spanish Disease, Symptom, and Procedure annotations into six languages. Evaluation used MultiClinAI gold standard with strict span matching and character-overlap F1 Results: Direct LLM projection achieved the strongest and most consistent performance. GLM 5.2 obtained a mean Strict F1 of 0.9201 across 18 language-entity combinations, while locally deployable Gemma4:31B achieved 0.9133. The best LLM configuration improved Strict F1 over the previous state of the art in all 18 settings, by 0.0564-0.1512, yielding 55,416 grounded mentions with reconstructed offsets. Conclusions: Direct LLM-based projection enables high-quality multilingual clinical annotation transfer and provides a practical approach for extending clinical NLP resources to languages with fewer annotated datasets and language-specific tools. Combined with local inference and deterministic validation, it can substantially reduce expert time and cost for multilingual clinical corpus construction.
发表机构
- Universidad de Málaga(马拉加大学)
- Research Institute of Multilingual Language Technologies, Universidad de Málaga(马拉加大学多语言语言技术研究所)
- IBIMA Plataforma BIONAND, Instituto de Investigación Biomédica de Málaga(IBIMA BIONAND平台,马拉加生物医学研究所)
机构由 AI 辅助整理,请以论文原文为准。