arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

缩小英阿医学知识鸿沟:通过因果层选择实现针对性低秩适配

Bridging the English-Arabic Medical Knowledge Gap: Targeted Low-Rank Adaptation via Causal Layer Selection

Chaimae Abouzahir, Musa Khan, Hala Ali-Hassan, Congbo Ma, Khaled Saleh, Yousra Sadqi, Jihad Mallat, Walid Al-Eisawi, Nizar Habash, Farah E. Shamout

arXiv 2608.00207首次发表:更新:

发表机构

New York University Abu Dhabi; Cleveland Clinic Abu Dhabi(纽约大学阿布扎比分校; 克利夫兰诊所阿布扎比分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对阿拉伯语医学LLMs性能弱于英文的问题,提出TLoRA方法并构建阿拉伯医学对话基准,通过机制诊断实现针对性适配,提升了阿拉伯语医学任务性能。

AI 中文摘要

大型语言模型(LLMs)在英文医学任务中表现强劲,但在阿拉伯语任务中性能大幅下降,这一鸿沟被广泛归因于训练数据有限。我们通过调优透镜探测和因果激活修补系统研究了这一假设,发现阿拉伯语医学知识存在于模型中间表示中,但无法在输出层显现。这一机制洞察催生了针对性适配策略:我们提出针对性低秩适配(TLoRA),而非微调整个网络,仅将其限制在跨语言表示发生分歧的层窗口,即性能失效的输出层上游。我们在多项选择医学问答任务上评估了TLoRA,结果显示该方法优于全网络LoRA、零样本和少样本基线。我们还在短答案生成和多轮临床对话任务上进行评估,其表现具有竞争力且无需特定任务微调。此外,我们引入了AraClinicDialog,这是由临床医生构建的现代标准阿拉伯语(MSA)医学对话基准,包含经验证的四种阿拉伯语变体。这些贡献共同表明,机制诊断可作为代表性不足语言医学大型语言模型针对性适配的实用指南。

英文摘要

Large Language Models (LLMs) perform strongly in English medical tasks but degrade substantially in Arabic, a gap widely attributed to limited training data. We systematically investigate this assumption via tuned lens probing and causal activation patching, and find that Arabic medical knowledge is present in intermediate model representations but fails to surface at the output. This mechanistic insight motivates a targeted adaptation strategy: rather than fine-tuning the full network, we propose Targeted Low-Rank Adaptation (TLoRA), restricted to the layer window where cross-lingual representations diverge, upstream of the output layers where the failure manifests. We evaluate TLoRA on multiple-choice medical QA, where our approach outperforms full-network LoRA, zero-shot, and few-shot baselines. We further evaluate it on short-answer generation and multi-turn clinical dialogue, where it performs competitively without the need for task-specific finetuning. We additionally introduce AraClinicDialog, a clinician-constructed Arabic medical dialogue benchmark in MSA with validated variants across four Arabic dialects. Together, these contributions demonstrate that mechanistic diagnosis can serve as a practical guide for targeted adaptation in underrepresented-language medical LLMs.

Journal refFindings of the Association for Computational Linguistics: EMNLP 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑