发表机构
Johns Hopkins University; Merck & Co., Inc.(约翰霍普金斯大学; 默克公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究审计多智能体假设生成中的成对等价性判断,发现自我批判轮次增加机制级分歧,且等价规则(LLM评判者、TF-IDF、嵌入)显著影响多样性测量,表明等价性定义是关键的测量选择。
AI 中文摘要
基于大语言模型(LLMs)的多智能体系统日益应用于科学发现和假设生成。在实验事实依据存在之前,细化效应和所交付集合的多样性都难以解释,且两者通常通过判断生成的假设对是否描述相同的底层机制来报告。我们研究两个依赖于这种成对等价性判断的评估问题:(1)自我批判在多大程度上改变了交付的假设,超出运行间变异性;(2)用于对假设进行分组的等价规则如何影响所测量的多样性。在四个专有实例中,我们保持初始假设固定,以0、1和5轮批判重新运行下游工作流,并使用LLM作为评判者对匹配的假设对进行评分。相对于匹配的相同深度重运行,从0轮到1轮产生了34.5个百分点(pp)的额外机制级分歧,而从1轮到5轮增加了1.3个百分点。然后,我们比较了三种等价规则:词频-逆文档频率(TF-IDF)相似度、密集嵌入和相同的LLM评判者。我们构建了受控的假设对,这些假设对要么通过措辞或生物学术语变化保留因果解释,要么在保持其余部分固定的同时替换因果链的一个组成部分。所有三种规则对保持意义的编辑均不变,但当初始事件被替换时,LLM评判者将83%的有效对识别为不同机制,而TF-IDF为0%,嵌入为8%;仅改变定义相同机制的评分标准,这一数字从38%变为96%。综合来看,这些结果表明成对等价性判断是一种测量选择:机制等价性的定义方式既影响自我批判效应的估计,也影响所生成假设的测量多样性。
英文摘要
Multi-agent systems built on large language models (LLMs) are increasingly applied to scientific discovery and hypothesis generation. Both the effect of refinement and the diversity of the delivered set are hard to interpret before experimental ground truth exists, and both are typically reported by deciding whether pairs of generated hypotheses describe the same underlying mechanism. We study two evaluation questions that rest on this pairwise equivalence judgment: (1) how much self-critique changes delivered hypotheses beyond run-to-run variability, and (2) how the equivalence rule used to group hypotheses affects measured diversity. Across four proprietary instances, we hold opening hypotheses fixed, rerun the downstream workflow with 0, 1, and 5 critique rounds, and score matched hypothesis pairs with an LLM-as-a-judge. Relative to matched same-depth reruns, moving from 0 to 1 round produces 34.5 percentage points (pp) of additional mechanism-level divergence, whereas 1 to 5 rounds adds 1.3 pp. We then compare three equivalence rules: term frequency--inverse document frequency (TF--IDF) similarity, dense embeddings, and the same LLM-as-a-judge. We construct controlled hypothesis pairs that either preserve the causal explanation through wording or biological-terminology changes, or replace one component of the causal chain while holding the rest fixed. All three rules are invariant to meaning-preserving edits, but when the initiating event is replaced, the LLM-as-a-judge identifies 83% of valid pairs as different mechanisms, versus 0% for TF--IDF and 8% for embeddings; varying only the rubric that defines same mechanism moves this figure from 38% to 96%. Together, these results show that pairwise equivalence judgments are a measurement choice: how mechanism equivalence is defined affects both the estimated effect of self-critique and the measured diversity of generated hypotheses.
CommentsAccepted at the Agentic AI for Biological Discovery Workshop at NeurIPS 2026