AI 中文总结
研究针对大语言模型科学发现智能体首个研究问题难审计的问题,提出FirstResearch框架,核心是研究问题证书,记录多种要素使问题可检查。实验表明该框架在多个主题上优于基线,以证书为中心的核心组件作用显著,支持明确推导约束可提升问题可审计性。
AI 中文摘要
用于科学发现的大语言模型系统越来越多地协助构思、文献综合、实验规划和报告生成,但它们提出的第一个研究问题可能仍难以审计:它可能听起来合理,但却不暴露科学家应检查的机制、证伪者或假设。我们引入了FirstResearch,这是一个用于科学大语言模型智能体的第一性原理研究问题形成框架,其核心产物是一个结构化的研究问题证书。该证书记录了原始定义、假设、机制模型、矛盾或冲突、可证伪假设、最小决定性测试和失败更新规则,使得所提出的问题在下游执行之前是可检查的。在十个大语言模型智能体研究主题上,在主要的DeepSeek盲评协议下,FirstResearch优于受AI共同科学家、智能体实验室和AI Scientist-v2启发的受控提示级基线。对相同的40个基线包进行Gemini-2.5-Flash独立评判重新评分保留了系统级排名,FirstResearch得分为4.86/5,最强基线得分为4.38/5,平均得分的皮尔逊一致性为0.865。一个单重复消融检查点进一步表明,以证书为中心的核心是最强的组件:仅证书评分在DeepSeek下达到4.90/5,在Gemini下达到4.88/5,而在两位评判下移除证书得分均低于1/5。这些结果是初步的,使用大语言模型评判而非人类领域专家,但它们支持一个狭义的科学发现主张:明确的推导约束是使大语言模型生成的科学问题更具可审计性的有前景机制。代码、提示、保存的输出和重现脚本可在这个https URL获取。
英文摘要
LLM systems for scientific discovery increasingly assist with ideation, literature synthesis, experiment planning, and report generation, but the first research question they propose can remain difficult to audit: it may sound plausible without exposing the mechanism, falsifier, or assumption that a scientist should inspect. We introduce FirstResearch, a first-principles research-question formation framework for scientific LLM agents whose core artifact is a structured Research Question Certificate. The certificate records primitive definitions, assumptions, a mechanism model, a tension or contradiction, a falsifiable hypothesis, a minimal decisive test, and a failure update rule, making the proposed question inspectable before downstream execution. On ten LLM-agent research topics, FirstResearch outperforms controlled prompt-level baselines inspired by AI co-scientist, Agent Laboratory, and AI Scientist-v2 under a primary DeepSeek-blind-judge protocol. A Gemini-2.5-Flash independent-judge rescore of the same 40 baseline packages preserves the system-level ranking, with FirstResearch scoring 4.86/5 versus 4.38/5 for the strongest baseline and Pearson agreement of 0.865 on average score. A one-repeat ablation checkpoint further suggests that the certificate-centered core is the strongest component: certificate-only scoring reaches 4.90/5 under DeepSeek and 4.88/5 under Gemini, while removing certificates drops below 1/5 under both judges. These results are preliminary and use LLM judges rather than human domain experts, but they support a narrow scientific-discovery claim: explicit derivation constraints are a promising mechanism for making LLM-generated scientific questions more auditable. Code, prompts, saved outputs, and reproduction scripts are available at https://github.com/louiswang524/FirstResearch.
CommentsWithdraw for further improvement and work consolidation