AI 中文总结
研究聚焦自动化研究系统隐性失败问题,核心方法是提出ARA框架,集成多环节到统一管道,将自然语言问题转化为可执行代码并评估。主要贡献是改变失败模式,强调有效性优先的自动化科学系统评估应综合考量答案准确性及因果声明合理性。
AI 中文摘要
自动化研究系统虽有望加速实证分析,但易出现隐性失败,即分析代码成功执行却依赖无效因果假设。我们提出基于人工智能的流行病学研究助手(ARA)框架,通过明确编码因果设计原则、特定研究假设和方法约束,使这些失败可见。ARA将协议构建、合成数据生成和对抗验证集成到统一管道中。它先构建协议,再用具有已知真实效应的结构因果模型生成合成数据集,将自然语言研究问题转化为结构化因果协议和可执行分析代码。生成的分析在识别假设的受控违反情况下进行评估。我们在自动因果推理基准上评估ARA,结果表明协议构建和对抗验证虽未始终改善与基准估计的数值一致性,但改变了失败模式,凸显了协议问题等。这表明有效性优先的自动化科学系统不仅应按答案准确性评估,还应看其是否能指出因果声明何时不合理。
英文摘要
While automated research systems promise to accelerate empirical analysis, they are prone to silent failures: instances in which analysis code executes successfully yet relies on invalid causal assumptions. We present the Artificial Intelligence (AI)-based Epidemiology Research Assistant (ARA), a framework that makes these failures visible by explicitly encoding causal design principles, study-specific assumptions, and methodological constraints. ARA integrates protocol construction, synthetic data generation, and adversarial validation into a unified pipeline. The framework translates natural language research questions into structured causal protocols and executable analysis code by first constructing a protocol and then generating synthetic datasets using Structural Causal Models (SCMs) with known ground-truth effects. This synthetic-data step can also support pipeline development when access to confidential data, such as medical data, is restricted. The generated analysis is then evaluated under controlled violations of identification assumptions. We evaluate ARA on the Automated Causal Reasoning Benchmark, assessing recovery of identification strategies, causal quantities, treatment and outcome variables, and consistency between generated code and approved protocol. Protocol construction and adversarial validation did not consistently improve numerical agreement with benchmark estimates compared with standard LLM-based generation. However, they changed the failure mode: instead of silently returning causal estimates, ARA often surfaced protocol concerns, diagnostic failures, incomplete inference, or downgraded non-causal interpretations. These findings suggest that validity-first automated science systems should be evaluated not only by answer accuracy, but also by whether they indicate when causal claims are unwarranted.