SGHA:基于证据的研究问题发现,使用本地语言模型
SGHA: Evidence-Grounded Research Problem Discovery with Local Language Models
- C-MInDS, IIT Bombay(印度理工学院孟买分校C-MInDS)
- Machine Learning Department, MBZUAI(穆罕默德·本·扎耶德人工智能大学机器学习系)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
该研究提出SGHA,一种基于本地9B开源LLM的研究问题发现系统,通过构建文献证据图检测研究差距,经对比实验证明其可实现可核查的研究问题构建,无需依赖专有前沿模型。
AI中文摘要:
近期实现完全自动化AI科学家的相关研究表明,语言模型智能体可生成假设、执行实验并撰写科学手稿。但在研究初期的研究问题构建阶段,这些AI科学家往往严重依赖专有前沿模型。其提案受不透明的参数知识及基于提案本身的文献搜索影响,此类知识如同黑箱,这种依赖使得生成研究问题的证据基础与有效性难以核查,且该过程易受模型特有的幻觉和偏差影响。此外,若专有研究材料被传输至外部API,使用这些模型会引发保密性、隐私及数据治理问题。我们提出Structural Gap Hypothesis Agent(SGHA,结构差距假设智能体),这是一个完全自动化、以语料库为核心的研究问题发现系统,可完全在本地语言模型(LLM)上运行。SGHA将科学文献语料库构建为证据关联的论文对象与类型化证据图,检测论文间未解决的结构模式,在构建前筛选候选差距,并生成可追溯的研究问题系列,尤其能输出假设、目标、成功标准及剩余模糊点。SGHA的所有基于LLM的组件均使用本地部署的9B参数开源权重语言模型执行,无需依赖专有前沿模型API。我们在五个机器学习领域将SGHA与AI Scientist-v2的想法构建模块进行对比,结果显示,明确的语料库结构与证据约束推理可支持有前景、可核查的研究问题构建,且在生成或验证阶段无需依赖前沿模型。
英文摘要:
Recent efforts toward fully automated AI scientists have demonstrated that language-model agents can generate hypotheses, execute experiments, and draft scientific manuscripts. However, during the early stages of research, when research problems are formulated, these AI scientists often rely heavily on proprietary frontier models. Their proposals are shaped by opaque parametric knowledge and by literature searches conditioned on the proposals themselves. Such knowledge is effectively a black box, and this dependence makes the evidential basis and validity of generated research problems difficult to audit and leaves the process vulnerable to model-specific hallucinations and biases. Furthermore, if proprietary research materials are transmitted to external APIs, the use of these models creates confidentiality, privacy, and data-governance concerns. We introduce the Structural Gap Hypothesis Agent (SGHA), a fully automated, corpus-first research-problem discovery system that runs entirely on a local LLM. SGHA structures a scientific literature corpus into evidence-linked paper objects and a typed evidence graph, detects unresolved structural patterns across papers, screens candidate gaps before formulation, and produces traceable research-problem families. In particular, it is able to output assumptions, objectives, success criteria, and remaining ambiguities. All LLM-based components of SGHA are executed using a locally served open-weight 9B language model, without requiring proprietary frontier-model APIs. We compare SGHA with the AI Scientist-v2 idea formulation module in five machine-learning domains. Our results suggest that explicit corpus structure and evidence-constrained reasoning can support promising, inspectable research-problem formulation without relying on frontier models during generation or verification.