AI 中文总结
研究针对现有在线医疗咨询基准与实际临床实践契合度差的问题,构建MedRealMM基准测试,用多模态临床挑战点提取框架转化任务并设评分标准,评估多个大语言模型,指出图像信息重要,前沿模型有不足,为多模态医学推理评估提供现实可重复基准。
AI 中文摘要
大语言模型越来越多地应用于在线医疗咨询,但现有基准与实际临床实践的契合度仍较差。许多基准依赖合成对话或患者模拟器,遗漏患者上传的医学图像,或使用不能很好反映临床质量的指标评估开放式临床反应。我们引入MedRealMM,这是一个基于全国性中文互联网医院收集的匿名医患互动构建的多模态在线医疗咨询大规模基准测试。MedRealMM使用多模态临床挑战点提取框架识别真实咨询轨迹中具有临床挑战性的时刻,并将其转化为标准化的下一轮回复生成任务,同时保留前文的文本-图像上下文。每个实例都配有由医生完善的特定案例评分标准,奖励符合临床要求的行为,惩罚不安全、无依据或矛盾的回复。当前版本包含跨越64个临床科室的5620个真实世界多模态案例。我们评估了19个通用和医学专用大语言模型,包括纯文本和多模态系统。结果表明图像信息对可靠的临床性能至关重要,当前前沿模型仍不及在线医生的回复。尽管一些前沿模型满足的积极临床标准与医生一样多甚至更多,但它们触发的消极标准更多,这表明避免安全敏感错误仍是核心瓶颈。MedRealMM为评估真实世界在线咨询中的多模态医学推理提供了一个现实且可重复的基准测试。该数据集将在Hugging Face上公开提供。
英文摘要
Large language models (LLMs) are increasingly deployed in online medical consultation, yet existing benchmarks remain poorly aligned with real clinical practice. Many rely on synthetic conversations or patient simulators, omit patient-uploaded medical images, or evaluate open-ended clinical responses using multiple-choice or lexical-overlap metrics that poorly reflect clinical quality. We introduce \textbf{MedRealMM}, a large-scale benchmark for multimodal online medical consultation built from de-identified patient-doctor interactions collected from a nationwide Chinese internet hospital. MedRealMM uses a Multimodal Clinical Challenge Point (MCCP) extraction framework to identify clinically demanding moments in authentic consultation trajectories and converts each into a standardized next-response generation task while preserving the preceding text-image context. Each instance is paired with a case-specific rubric refined by physicians that rewards clinically desirable behaviors and penalizes unsafe, unsupported, or contradictory responses. The current release contains 5,620 real-world multimodal cases spanning 64 clinical departments. We evaluate 19 general-purpose and medical-specialized LLMs, including text-only and multimodal systems. Our results show that image information is critical for reliable clinical performance and that current frontier models remain below the online physician response. Although some frontier models satisfy as many or more positive clinical criteria than physicians, they trigger more negative criteria, indicating that safety-sensitive error avoidance remains a central bottleneck. MedRealMM offers a realistic and reproducible benchmark for evaluating multimodal medical reasoning in real-world online consultation. The dataset will be publicly available on Hugging Face at https://huggingface.co/datasets/jdh-algo/MedRealMM.