发表机构
Looloo Health; Kasetsart University; Mahidol University(卢卢健康; 泰国农业大学; 玛希隆大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出ASCRIBE框架,通过为对话中原子事实赋予临床显著性等级提升泰语SOAP笔记生成的完整性与可靠性,并发布首个泰语临床基准ThaiClinicBench及合成语料,显著优于现有提示方法。
AI 中文摘要
自动生成SOAP笔记可减轻医生的文书负担,但现有推理方法常遗漏临床重要信息并生成缺乏依据的内容。泰语领域的进展进一步受到缺乏公开数据集的阻碍。我们提出ASCRIBE,一种受医生启发的推理框架,在摘要前为对话中提取的每个原子事实赋予临床显著性等级,使通用大语言模型成为更可靠的记录员。我们还发布了ThaiClinicBench,这是首个基于真实就诊的去标识化泰语临床摘要基准,同时提供了由真实临床笔记衍生的合成训练语料。作为提示词,ASCRIBE在GPT-5.4和Gemini 3.1 Pro上,在医生对齐的LLM评判指标上优于思维链提示,并在完整性LLM评判指标上比标准提示提升最多10.3个百分点。作为GRPO奖励,它使仅基于合成数据训练的Gemma-4-E4B模型在事实精确度上匹配Gemini 3.1 Pro,并在完整性上超越之。代码和数据可在该https URL获取。
英文摘要
Automatic SOAP note generation can ease the documentation burden on physicians, but existing reasoning methods often omit clinically important information and generate unsupported content. Progress in Thai is further hindered by the lack of publicly available datasets. We propose ASCRIBE, a physician-inspired reasoning framework that ascribes a clinical-significance level to each extracted atomic fact in the conversation before summarization, making a general-purpose LLM a more reliable scribe. We also release ThaiClinicBench, the first de-identified Thai clinical summarization benchmark of real encounters, together with a synthetic training corpus derived from real clinical notes. As a prompt, ASCRIBE outperforms chain-of-thought prompting on GPT-5.4 and Gemini 3.1 Pro across the physician-aligned LLM-judge metrics and improves on standard prompting by up to 10.3 points on the completeness LLM-judge metric. As a GRPO reward, it enables a Gemma-4-E4B model trained solely on synthetic data to match Gemini 3.1 Pro in factual precision and surpass it in completeness. Code and data can be found at https://github.com/loolootech/ascribe.