arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.03635cs.AI

用于药物毒性预测的提示工程分析

Analysis of Prompt Engineering for Drug Toxicity Prediction

Mia MacGregor, Aakash Welgamage Don, Mark Bartlett

首次发表
浏览论文内容

中文总结 AI 辅助

本文分析用于药物毒性预测的提示工程,发现LLMs的自然差异超过提示微调,但用化学信息学代码提取特征可显著提升模型性能,该方法适用于生物信息学多领域。

中文摘要 AI 辅助

英国的临床试验成本高达130万英镑,药物失败率约为90%,毒性是药物失败的主要影响因素,相关测试既耗时又成本高昂。近年来,人们越来越多地探索利用人工智能辅助药物毒性预测,其中大型语言模型(LLMs)的应用十分广泛。然而,当对提示进行微小修改时,LLMs会表现出显著的差异,这引发了人们对其提示工程敏感性的担忧。提示工程用于优化提供给LLMs的提示,以生成期望的输出。本文提出了一种分析药物毒性预测提示工程的方法,旨在研究提示措辞对药物毒性预测的重要性。研究人员向LLMs发出提示,以识别预测药物毒性时具有重要意义的化学性质,并构建了提示来探究角色、提示结构和规则解释。随后,利用LLMs从初始提示中识别出的特征生成数据集,再将该数据集输入机器学习算法。实验结果表明,LLMs中出现的自然差异超过了对提示进行的任何微调;不过,当使用化学信息学代码提取特征而非使用LLMs生成的值时,模型性能会有显著提升。所提出的分析方法适用于生物信息学不同领域的多种提示类型。

英文摘要

Clinical trials in the UK can cost up to £1.3 million, with approximately 90% drug failure rate. Toxicity is a major contributing factor in drug failure. Testing is time and cost intensive. In recent years, the use of artificial intelligence has been increasingly explored to aid in the prediction of drug toxicity, with extensive use of large language models (LLMs). However, LLMs can show considerable variation when minor changes are made to prompts, which raises concerns about their sensitivity to prompt engineering. Prompt engineering is used to optimise a prompt given to an LLM to generate the desired output. This paper proposes a method to analyse prompt engineering for drug toxicity prediction. The aim of the paper is to investigate the importance of prompt phrasing for drug toxicity prediction. LLMs were prompted to identify chemical properties of significance when predicting drug toxicity. Prompts were constructed to investigate; job role, prompt structuring, and rule interpretation. LLMs were then used to generate datasets, using the identified features from initial prompting, which were then passed to machine learning algorithms. The experiments show that the natural variance which occurs in LLMs outweighs any fine-tuning of prompts. There were, however, substantial improvements in model performance when using chemoinformatic code to extract features instead of using LLM-generated values. The proposed analysis methodology is applicable to a wide range of prompt types across different areas of bioinformatics.

发表机构

  • Robert Gordon University(罗伯特戈登大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑