通过人类对齐设计用于药物发现中智能体AI的鲁棒性大语言模型评估系统
Designing a Robust LLM-Based Evaluation System for Agentic AI in Drug Discovery Through Human Alignment
浏览论文内容
中文总结 AI 辅助
本研究针对药物发现智能体AI,提出人类对齐的LLM作为评判者评估框架,优化后对齐度达0.86,为科学领域智能体系统评估提供可复用模板。
中文摘要 AI 辅助
智能体大语言模型(LLM)系统正在重塑化学与药物发现领域的科研工作流程,但对其开放式、工具增强型输出的评估仍是核心瓶颈。BLEU、ROUGE等基于参考的指标无法捕捉语义正确性,而专家人工评估无法匹配这些系统所需的迭代速度。LLM作为评判者(LLM-as-a-Judge)范式已成为一种可扩展的替代方案,但现有的药物发现基准部署LLM评判者时未验证其与人类专家的对齐性。本研究针对阿斯利康部署的智能体药物发现助手ChatInvent,提出一种LLM作为评判者的评估框架,包含四项贡献:第一,定义了四项输出质量评估维度——完整性、相关性、结构清晰度与范围一致性,同时辅以确定性工具调用正确性检查;第二,通过与五名专家标注者开展的人类对齐研究验证该评判者,对比Gemini 3.1 Pro、Claude Opus 4.7、GPT-5与Llama 3.1 70B作为候选评判者;第三,使用人工标注示例的少样本演示优化表现最佳的评判者,将其与人类多数投票的对齐度从0.80提升至0.86;第四,将优化后的评判者应用于70个保留问题,揭示具体局限性并发现非正式措辞不会系统性降低输出质量,相反,让LLM在查询智能体重写原始问题是有益的。本框架为科学领域智能体系统的人类对齐评估提供了可复用模板。
英文摘要
Agentic large language model (LLM) systems are reshaping scientific workflows in chemistry and drug discovery, but evaluating their open-ended, tool-augmented outputs remains a fundamental bottleneck. The LLM-as-a-Judge paradigm has emerged as a scalable alternative, but existing drug discovery benchmarks deploy LLM judges without validating their alignment with human experts. In this work, we present an LLM-as-a-Judge evaluation framework for ChatInvent, an agentic drug discovery assistant deployed at AstraZeneca, with five contributions. First, we define four output-quality evaluation dimensions---Completeness, Relevancy, Structural Clarity, and Scope Adherence---alongside deterministic Tool Call Correctness checks. Second, we validate the judge through a human alignment study with five expert annotators, comparing Gemini 3.1 Pro, Claude Opus 4.7, GPT-5, and Llama 3.1 70B as candidate judges. Third, we optimize the best-performing judge using few-shot demonstrations of human-annotated examples, improving alignment with the human majority vote from 0.80 to 0.86. Fourth, applying the optimized judge to 70 held-out questions, we surface concrete limitations and find no strong evidence that informal phrasing degrades output quality; it may, however, still be helpful to have the LLM rewrite the original question before querying the agent. Finally, we extend the framework to 38 adversarial questions that are ambiguous, invalid, out-of-scope or ethically sensitive, and show that the agent's refusal behavior is guided by the stated intent of a request. Our framework provides a reusable template for human-aligned evaluation of agentic systems in scientific domains.
发表机构
- AstraZeneca(阿斯利康)
- Chalmers University of Technology(查尔姆斯理工大学)
- University of Gothenburg(哥德堡大学)
- Science for Life Laboratory (SciLifeLab)(生命科学实验室)
机构由 AI 辅助整理,请以论文原文为准。