AI 中文总结
本研究采用双重评估框架,结合人类专家与LLM-as-a-judge,评估癌症患者护理应用中AI生成摘要的质量,发现其存在遗漏和轻微不准确问题,以此优化相关设计与防护措施。
AI 中文摘要
大型语言模型(LLMs)正越来越多地被集成到数字健康平台中,用于生成复杂医疗数据的摘要。尽管这些模型能够提升患者参与度与沟通效果,但此类系统也引发了临床场景下的准确性、忠实性及安全性方面的担忧。本研究在癌症患者护理应用中,采用双重评估框架对AI生成的摘要进行评估:肿瘤临床医生、面向患者的护理人员等人类领域专家,从准确性、临床相关性、可读性维度提供了摘要质量的真实评估;同时,我们将LLMs作为评估者(LLM-as-a-judge)开展并行评估。研究发现生成摘要存在部分局限性,例如偶尔出现内容遗漏和轻微不准确问题,我们对这些问题进行了系统分析,并用于迭代优化提示设计、内容依据及安全防护措施。
英文摘要
Large language models (LLMs) are increasingly being integrated into digital health platforms to generate summaries of complex medical data. Although these models can improve patient engagement and communication, these systems also raise concerns about accuracy, faithfulness, and safety in clinical contexts. In this study, we evaluate AI-generated summaries within a cancer patient care application using a dual assessment framework. Human domain experts, including oncology clinicians and patient-facing care staff, provided ground-truth evaluations of summary quality along dimensions of accuracy, clinical relevance, and readability. In parallel, we employed LLMs serving as evaluators (LLM-as-a-judge). Some limitations were identified in the generated summaries e.g., occasional omissions and minor inaccuracies. These were systematically analyzed and used to iteratively improve prompt design, grounding, and safety guardrails.