arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.06022cs.CLq-bio.GN

EpiBench:大型语言模型能否为抗体药物发现表位?

EpiBench: Can LLMs Understand Epitopes for Antibody Drug Discovery?

  • Zhejiang University(浙江大学)

机构由 AI 辅助整理,请以论文原文为准。

Zirui Wang, Jiaqi Wang, Qinghan Wang, Yuzhi Xu, Gang Du, Tingjun Hou, Odin Zhang

AI总结:

本研究提出用于评估大型语言模型表位推理能力的基准EpiBench,发现现有通用LLMs在抗体药物发现相关表位推理上仍存局限,为改进序列感知型生物医学LLMs提供了测试平台。

AI中文摘要:

表位决定抗体结合抗原的位置,并塑造下游治疗特性,如功能阻断和逃逸抗性,因此表位理解是抗体药物发现的核心。尽管大型语言模型(LLMs)展现出强大的生物医学推理能力,但目前仍不清楚它们能否直接从抗原和抗体序列中推断表位信息。现有表位资源通常聚焦于孤立的预测任务或依赖特定结构设置,而通用蛋白质基准未在抗体开发全流程中评估以表位为中心的决策。为解决这一缺口,我们提出EpiBench——一种闭卷、基于序列且可自动评分的基准,用于评估LLMs的表位推理能力。EpiBench包含1609个经整理的样本,这些样本基于抗体-抗原接触结构、经整理的功能性B细胞测定以及深度突变扫描逃逸测量构建,涵盖五个关联任务:可靶向区域发现、抗体条件下表位识别、表位分箱、功能性表位评估以及抗体逃逸评估,且通过控制采样减少基于捷径的评估伪影。我们评估了9个通用LLMs,并通过任务特定基线、抗原长度分层、显式推理对比及失败模式检查分析其行为。结果显示,当前LLMs能捕捉部分表位相关信号,但在抗体特定序列定位、长上下文残基定位及基于生物学的推理方面仍存在局限。因此,EpiBench为测量和改进序列感知型生物医学LLMs提供了诊断测试平台,助力实现可靠的LLM辅助抗体发现。

英文摘要:

Epitopes determine where antibodies bind antigens and shape downstream therapeutic properties such as functional blockade and escape resistance, making epitope understanding central to antibody drug discovery. Although large language models (LLMs) have shown strong biomedical reasoning ability, it remains unclear whether they can infer epitope information directly from antigen and antibody sequences. Existing epitope resources typically focus on isolated prediction tasks or rely on specialized structural settings, while general protein benchmarks do not evaluate epitope-centered decisions across the antibody development workflow. To address this gap, we introduce EpiBench, a closed-book, sequence-based, and automatically scorable benchmark for evaluating epitope reasoning in LLMs. EpiBench contains 1,609 curated samples grounded in structural antibody--antigen contacts, curated functional B-cell assays, and deep mutational scanning escape measurements. It covers five connected tasks: targetable region discovery, antibody-conditioned epitope identification, epitope binning, functional epitope assessment, and antibody escape assessment, with controlled sampling to reduce shortcut-based evaluation artifacts. We evaluate nine general-purpose LLMs and analyze their behavior through task-specific baselines, antigen length stratification, explicit-reasoning comparison, and failure-mode inspection. The results show that current LLMs capture partial epitope-related signals but remain limited in antibody-specific sequence grounding, long-context residue localization, and biologically grounded reasoning. Therefore, EpiBench provides a diagnostic testbed for measuring and improving sequence-aware biomedical LLMs toward reliable LLM-assisted antibody discovery.

↑