arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.02949cs.CL

通过多编程语言指令调优与集成方法增强生物医学命名实体识别

Enhancing Biomedical Named Entity Recognition via Multiple Programming Languages Instruction Tuning and Ensemble Method

Songtao Li, Yijia Zhang, Jianyuan Yuan, Shidi Zhang, Fengyu Zhang, Hongfei Lin

首次发表
浏览论文内容

中文总结 AI 辅助

提出MITE方法,通过多编程语言指令调优和实体级投票集成,将BioNER转化为结构到结构生成任务,在六个数据集上优于现有基线。

中文摘要 AI 辅助

指令调优已成为将大型语言模型(LLMs)应用于生物医学命名实体识别(BioNER)的常见范式。然而,现有的指令调优方法仍面临两个关键挑战。首先,传统的自然语言指令通常将BioNER注释序列化为扁平文本输出,对类型化实体提取提供的结构约束有限。其次,高质量的生物医学注释有限,从单一序列化输出形式学习可能限制结构多样性并降低模型鲁棒性。尽管可以引入外部生物医学知识来缓解数据稀缺问题,但这通常需要昂贵的资源构建。为解决这些挑战,我们提出了MITE,一种用于BioNER的多编程语言指令调优与集成方法。MITE通过将指令和实体输出表示为代码格式表示,将BioNER重新表述为结构到结构的生成任务。具体而言,每个训练实例被转换为多种编程语言格式,包括Python、C++和Java,同时保持相同的底层实体语义。这些特定语言的表示提供了结构多样的监督,无需外部生物医学知识或额外注释。在推理过程中,MITE通过实体级投票策略聚合来自不同代码格式的预测,减少特定语言的预测方差并提高鲁棒性。在六个广泛使用的BioNER数据集上的实验表明,MITE始终优于代表性的基于BERT和基于LLM的基线,并展现出强大的跨数据集泛化能力。消融和参数分析进一步验证了所提出组件的有效性和鲁棒性。

英文摘要

Instruction tuning has become a common paradigm for applying large language models (LLMs) to biomedical named entity recognition (BioNER). However, existing instruction-tuning approaches still face two key challenges. First, conventional natural-language instructions typically serialize BioNER annotations as flat textual outputs, providing limited structural constraints for typed entity extraction. Second, high-quality biomedical annotations are limited, and learning from a single serialized output form may restrict structural diversity and reduce model robustness. Although external biomedical knowledge can be introduced to alleviate data scarcity, it often requires costly resource construction. To address these challenges, we propose MITE, a Multiple Programming Languages Instruction Tuning and Ensemble method for BioNER. MITE reformulates BioNER as a structure-to-structure generation task by representing both instructions and entity outputs in code-formatted representations. Specifically, each training instance is transformed into multiple programming-language formats, including Python, C++, and Java, while preserving the same underlying entity semantics. These language-specific representations provide structurally diverse supervision without requiring external biomedical knowledge or additional annotations. During inference, MITE aggregates predictions from different code formats through an entity-level voting strategy, reducing language-specific prediction variance and improving robustness. Experiments on six widely used BioNER datasets demonstrate that MITE consistently outperforms representative BERT-based and LLM-based baselines and exhibits strong cross-dataset generalization. Ablation and parameter analyses further verify the effectiveness and robustness of the proposed components.

发表机构

  • Dalian Maritime University(大连海事大学)
  • Beijing Institute of Technology(北京理工大学)
  • Northeastern University(东北大学)
  • Dalian University of Technology(大连理工大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑