arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

大型语言模型能否可靠地标注生物测定元数据以提升数据就绪度?

Can LLMs Reliably Annotate Bioassay Metadata to Improve Data Readiness?

Laura van Weesep, Riccardo Tedoldi, Jens Sjölund, Hossein Azizpour, Susanne Winiwarter, Ola Engkvist, Jon Paul Janet, Samuel Genheden, Juan Viguera Diez

arXiv 2610.01616首次发表:更新:

发表机构

AstraZeneca; Uppsala University; KTH Royal Institute of Technology; Science for Life Laboratory; Chalmers University of Technology; University of Gothenburg(阿斯利康; 乌普萨拉大学; 瑞典皇家理工学院; 生命科学实验室; 查尔姆斯理工大学; 哥德堡大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究评估了LLMs在PubChem大规模生物测定元数据标注中的可靠性,发现覆盖率严重不足,但LLMs能高召回率预测格式与检测技术,并辅助审计,表明其可支持大规模元数据管理,但需人工审查。

AI 中文摘要

用于分子性质预测的基础模型的兴起要求高度的AI数据就绪度,包括可靠的元数据标注。然而,公共数据库和工业筛选数据库都存在测定注释缺失、不一致或混淆的问题。在本研究中,我们量化了PubChem中BioAssay本体论(BAO)测定格式和物理检测方法字段注释缺失的程度,并调查了开源和专有大型语言模型(LLMs)是否能够直接从测定文本中可靠地预测和审计元数据注释。在我们的评估中,我们发现PubChem约200万生物测定中的注释覆盖率严重稀疏,36%缺乏测定格式,89%缺乏BioAssay类型,>99.9%缺乏任何BAO映射的测定格式或检测技术术语。这凸显了自动化测试元数据管理的必要性。利用从PubChem和ChEMBL衍生的评估集,我们评估了七个开源和专有LLMs与现有银标签的一致性。对于生物化学和基于细胞的测定格式,召回率至少为0.96,检测技术也有类似模式,尽管在代表性不足的类别上分歧增加。人工检查显示,许多分歧可追溯到银来源之间的不一致,而非LLM错误。此外,在与一位资深工业管理员的定性研究中,LLM生成的证据促使专家修改了他们自己的一些标签,表明LLMs可以标记可能被错误标注的测定。在整个研究中,专有和开源模型之间的性能差异很小。综合来看,这些结果表明LLMs可以支持大规模测定元数据的标注和审计,但在这些标签进入下游机器学习流水线之前,仍需按类别的可靠性估计和针对性的人工审查。

英文摘要

The emergence of foundation models for molecular property prediction requires a high degree of AI data readiness, including reliable metadata annotation. However, both public repositories and industrial screening databases suffer from missing, inconsistent, or conflated assay annotations. In this work, we quantify the extent of missing annotations in PubChem for the BioAssay Ontology (BAO) assay format and physical detection method fields and investigate whether open-source and proprietary large language models (LLMs) can reliably predict and audit metadata annotations directly from the assay text. In our assessment, we found that the annotation coverage across PubChem's $\sim$2 million bioassays is critically sparse, 36\% lacking an assay format, 89\% a BioAssay type, and >99.9\% any BAO-mapped assay format or detection technology term. This motivates the need for automated test-metadata curation. Using evaluation sets derived from PubChem and ChEMBL, we assess the agreement of seven open-source and proprietary LLMs with existing silver labels. Recall is at least 0.96 for biochemical and cell-based assay formats, with a similar pattern for detection technology, although disagreements increase on under-represented classes. Manual inspection shows that many of these disagreements trace back to inconsistencies between silver sources rather than to LLM error. Moreover, in a qualitative study with a senior industrial curator, LLM-generated evidence prompted the expert to revise some of their own labels, showing LLMs can flag potentially mislabeled assays. Across the study, performance differences between proprietary and open-source models were small. Together, these results suggest LLMs can support the large-scale annotation and auditing of assay metadata, though per-class reliability estimates and targeted human review remain necessary before such labels enter downstream ML pipelines.

CommentsAccepted to the AIDaR workshop at NeurIPS

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑