DISCERN:AI智能体能否像科学家一样工作并引导发现?
DISCERN: Can AI Agents Work Like Scientists and Guide Discovery?
- University of California San Diego(加州大学圣地亚哥分校)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
DISCERN基准通过三个层级测试AI智能体的数据完整性、分析验证和假设生成,发现现有模型在对抗性评审下表现不佳,未能实现可靠的自主科学发现。
AI中文摘要:
可靠的自动化研究要求智能体审查数据、验证分析,并基于可信证据生成假设,这有可能减少常规科学工作量,同时让科学家专注于解释和发现。现有基准通常仅分别评估分析任务完成或假设生成,而非测试可靠证据是否支持有效且新颖的论断。我们引入DISCERN(数据完整性与科学能力:证据、推理与新颖性),这是一个基于真实公开数据集的受控基准,评估自动化研究工作流程的三个关键层级。前两个层级测试数据完整性和在混杂因素及工具陷阱下的分析验证,第三个层级测试在对抗性评审下的假设生成与修订,包括反事实案例,其中与真实数据和已记录科学现象一致的证据与既定预期相冲突,从而激发替代解释和可检验假设。在203个任务、8个生命科学轨道和8个模型中,DISCERN表明,强大的总体表现可能掩盖层级特定的弱点。智能体在层级1中仅60.8%、层级2中34.2%、层级3中0.6%的评估中获得满分,惩罚归因于拒绝合理数据、未能将已识别的局限性纳入结论,以及假设生成中的广泛差异。按令牌和代码使用量的跨轨道排名比按证据判断的排名更稳定,表明计算努力的一致性高于基于证据的推理。这些概况识别了监督式科学辅助的机会,但当前智能体尚未展示可靠的自主分析或发现。代码和数据:此https URL
英文摘要:
Reliable automated research requires agents to vet data, verify analyses, and generate hypotheses grounded in trustworthy evidence, potentially reducing routine scientific workload while allowing scientists to focus on interpretation and discovery. Existing benchmarks often only assess analytical task completion or hypothesis generation separately rather than testing whether reliable evidence supports valid and novel claims. We introduce DISCERN (Data Integrity and Scientific Capability: Evidence, Reasoning, and Novelty), a controlled benchmark on real, publicly available datasets that evaluates three key levels of an automated research workflow. The first two levels test data integrity and analysis verification under confounds and tool traps, while the third tests hypothesis generation and revision under adversarial review, including counterfactual cases in which evidence consistent with real data and documented scientific phenomena conflicts with established expectations, motivating alternative explanations and testable hypotheses. Across 203 tasks, eight life-science tracks, and eight models, DISCERN shows that strong aggregate performance can mask level-specific weaknesses. Agents earn perfect scores in only 60.8% of Level 1, 34.2% of Level 2, and 0.6% of Level 3 evaluations, with penalties attributed to rejection of sound data, failure to carry recognized limitations into conclusions, and wide variation in hypothesis production. Cross-track rankings by token and code use are substantially more stable than rankings by evidence judgment, suggesting greater consistency in computational effort than in evidence-based reasoning. These profiles identify opportunities for supervised scientific assistance, but current agents do not yet demonstrate reliable autonomous analysis or discovery. Code and data: https://huggingface.co/datasets/discern-bench-anon/discern-benchmark