发表机构
Bangladesh University of Engineering and Technology (BUET)(孟加拉工程技术大学(BUET))
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对低资源语言孟加拉语医疗对话数据稀缺问题,构建了DocTalkBN多模态数据集,含大量真实远程医疗对话数据,并设置三项下游任务进行基准测试,为医疗NLP研究提供实用资源。
AI 中文摘要
可靠的医疗对话AI需要真实的医患交互数据,但这类数据集依然稀缺,尤其是对于孟加拉语等低资源语言。我们提出DocTalkBN,这是一个大规模多模态数据集,包含真实世界中孟加拉语专家远程医疗对话,数据来自全国播出的远程医疗节目,参与人员为拥有委员会认证的医生。DocTalkBN包含557.63小时的配对音频和文本、1515个多轮患者通话、10274个主持人-医生问答交互,总计170万 tokens,覆盖26个医学专科。与之前来自医疗论坛、书面健康内容或合成数据的资源不同,我们的数据集保留了低资源环境下真实医疗交互的自发性、上下文丰富性和口语特征。为支持基准驱动研究,我们进一步从语料库中构建了三个下游任务:医疗分诊分类、建议安全性评估和医疗命名实体识别,并对多种大型语言模型和基于编码器的基线进行了基准测试。我们的结果表明,DocTalkBN是一种实用资源,尤其适用于基于临床的推理任务。我们发布该资源,以促进未来关于可靠医疗NLP以及面向低资源语言更安全、更具文化根基的医疗保健系统的研究。我们的源代码和数据集可在此httpsURL获取。
英文摘要
Reliable medical conversational AI requires authentic expert--patient interaction data, yet such datasets remain scarce, especially for low-resource languages such as Bengali. We present DocTalkBN, a large-scale multimodal dataset of real-world expert telemedicine conversations in Bengali, collected from nationally broadcast telemedicine programs featuring board-certified physicians. DocTalkBN contains 557.63 hours of paired audio and text, 1,515 multi-turn patient calls, 10,274 host--doctor question--answer exchanges, totaling 1.7M tokens, spanning 26 medical specialties. Unlike prior resources derived from medical forums, written health content, or synthetic data, our dataset preserves the spontaneity, contextual richness, and spoken characteristics of authentic medical interactions in a low-resource setting. To support benchmark-driven research, we further construct three downstream tasks from the corpus, medical triage classification, advice safety evaluation, and medical named entity recognition, and benchmark a diverse set of large language models and encoder-based baselines. Our results show that DocTalkBN is a practically useful resource, particularly for clinically grounded reasoning tasks. We release this resource to facilitate future research on reliable medical NLP and safer, more culturally grounded healthcare systems for low-resource languages. Our source codes and dataset are publicly available at https://anonymous.4open.science/r/doctalk.