arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Gaokerena:一个小型波斯语医疗语言模型家族

Gaokerena: A Small Persian Medical Language Model Family

Mehrdad Ghassabi, Hamidreza Baradaran Kashani, Pedram Rostami, Sadra Hakim, Zahra Kazemi, Amirhossein Poursina, Milad Tavakoli, Audrina Ebrahimi

arXiv 2608.00932首次发表:更新:

发表机构

University of Isfahan; University of Tehran; University of Windsor; Alzahra University; University of Texas at Dallas(伊斯法罕大学; 德黑兰大学; 温莎大学; 阿勒扎哈拉大学; 德克萨斯大学达拉斯分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对波斯语医疗语言模型研究不足的问题,该研究推出Gaokerena家族,包含Gaokerena-V和Gaokerena-R两款模型,在医疗问答基准上取得性能提升,还配备不确定性头,为本地化数字医疗提供基础但仍需完善。

AI 中文摘要

人工智能在医疗问答系统中的集成已快速发展,但相关研究主要集中在英语领域,波斯语等低资源语言的研究严重不足。为解决这一缺口,本文推出Gaokerena,一个专为消费级硬件部署优化的新型小型波斯语医疗语言模型家族。作为本地化数字医疗的基础步骤,我们首先推出Gaokerena-V,该模型通过在新整理的9000万token波斯语医疗语料库和2万条经专家审核的医生问答对上训练基线模型,在翻译版医疗MMLU基准上的性能从46.28%提升至49.31%。其次,考虑到临床推理的关键需求,我们开发了Gaokerena-R,通过将思维链(Chain-of-Thought)方法与两个新型AI反馈强化学习(RLAIF)框架集成,优化基于偏好的推理。尽管使用与Gaokerena-V相同的基线架构且数据集更小,Gaokerena-R仍取得了52.98%的更高基准分数。此外,两个模型均配备定制开发的不确定性头,仅基于内部隐藏状态预测模型对其响应的置信度。这些结果虽表明波斯语医疗语言建模和主动安全评估取得显著进展,但当前性能水平仍不足以直接用于临床,凸显在实际部署前需进一步研究稳健的知识获取和严格的安全验证。

英文摘要

The integration of artificial intelligence into medical question-answering systems has advanced rapidly; however, research remains predominantly focused on English, leaving low-resource languages like Persian significantly underserved. To address this gap, this paper introduces Gaokerena, a novel family of compact Persian medical language models optimized for deployment on consumer-grade hardware. As a foundational step toward localized digital healthcare, we first present Gaokerena-V, developed by training a baseline model on a strategically selected subset of a newly curated 90-million-token Persian medical corpus (approximately 54 million tokens) together with 20,000 expert-vetted physician Q&A pairs (approximately 3 million tokens), for a total of 57 million new tokens. This training improved performance on a translated medical MMLU benchmark from 46.64% to 49.31%. Second, recognizing the critical demands of clinical reasoning, we developed Gaokerena-R by integrating a Chain-of-Thought approach with two novel Reinforcement Learning with AI Feedback (RLAIF) frameworks to optimize preference-based reasoning. Despite utilizing the same baseline architecture and a smaller dataset than Gaokerena-V, Gaokerena-R achieved a superior benchmark score of 52.98%. Furthermore, both models are equipped with custom-developed uncertainty heads that predict the models confidence in its responses based solely on internal hidden states. While these results demonstrate significant progress in Persian medical language modeling and proactive safety estimation, current performance levels remain insufficient for direct clinical application, highlighting the necessity for further research into robust knowledge acquisition and rigorous safety verification prior to real-world deployment.

Comments37 pages, 9 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑