arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2509.11803cs.CL

从模糊语音到医疗洞察:在含噪患者叙述上对 LLMs 进行基准测试

From Fuzzy Speech to Medical Insight: Benchmarking LLMs on Noisy Patient Narratives

  • Afeka Academic College of Engineering(阿费卡工程学院)
  • Holon Institute of Technology(霍隆理工学院)

机构由 AI 辅助整理,请以论文原文为准。

Eden Mama, Liel Sheri, Yehudit Aperstein, Alexander Apartsin

更新

AI总结:

针对大型语言模型解读非正式且含噪患者叙述的挑战,研究提出合成数据集 Noisy Diagnostic Benchmark (NDB),通过微调评估 BERT 和 T5 等模型在现实语言条件下的诊断能力。

AI中文摘要:

大型语言模型(LLMs)在医疗保健领域的广泛应用引发了关于其解读患者生成叙述能力的关键问题,这些叙述通常是非正式、模糊且含噪的。现有基准通常依赖于干净、结构化的临床文本,对模型在现实条件下的性能提供的洞察有限。在这项工作中,我们提出了一个新颖的合成数据集,旨在模拟患者自我描述,其特征是具有不同级别的语言噪声、模糊语言和外行术语。我们的数据集包含带有真实诊断标注的临床一致场景,跨越一系列沟通清晰度以反映多样化的真实世界报告风格。使用该基准,我们微调并评估了多个最先进的模型(LLMs),包括基于 BERT 的模型和编码器-解码器 T5 模型。为了支持可复现性和未来研究,我们发布了 Noisy Diagnostic Benchmark (NDB),这是一个包含含噪、合成患者描述的结构化数据集,旨在压力测试和比较大型语言模型(LLMs)在现实语言条件下的诊断能力。我们已向社区提供该基准:https://github.com/lielsheri/PatientSignal

英文摘要:

The widespread adoption of large language models (LLMs) in healthcare raises critical questions about their ability to interpret patient-generated narratives, which are often informal, ambiguous, and noisy. Existing benchmarks typically rely on clean, structured clinical text, offering limited insight into model performance under realistic conditions. In this work, we present a novel synthetic dataset designed to simulate patient self-descriptions characterized by varying levels of linguistic noise, fuzzy language, and layperson terminology. Our dataset comprises clinically consistent scenarios annotated with ground-truth diagnoses, spanning a spectrum of communication clarity to reflect diverse real-world reporting styles. Using this benchmark, we fine-tune and evaluate several state-of-the-art models (LLMs), including BERT-based and encoder-decoder T5 models. To support reproducibility and future research, we release the Noisy Diagnostic Benchmark (NDB), a structured dataset of noisy, synthetic patient descriptions designed to stress-test and compare the diagnostic capabilities of large language models (LLMs) under realistic linguistic conditions. We made the benchmark available for the community: https://github.com/lielsheri/PatientSignal

补充信息

↑