发表机构
Universidad Autónoma de Madrid; Spanish National Research Council (CSIC)(马德里自治大学; 西班牙国家研究委员会(CSIC))
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对气候政策双重差分研究,提出ARGUS语言模型管道,依据十一维度标准审计识别假设证据,必要时弃权,能检测73%缺陷,提供证据关联风险报告。
AI 中文摘要
双重差分(DID)研究被广泛用于评估气候政策,但评估支持其识别假设的证据仍然具有挑战性。我们引入了ARGUS,一个结构化的语言模型管道,它根据一个十一维度的假设-含义-证据评分标准来审计所报告的证据,并在无法检索到相关证据时弃权(不执行)。我们使用注入的缺陷、经济学论文和一个带有调和标签的小型试点来评估ARGUS。在11缺陷基准上,ARGUS检测到73%的植入缺陷,而基于关键词的管道仅为18%。在26篇经济学论文中,ARGUS在约40%的论文-维度评估上因缺乏可检索证据而弃权(不执行)。在一个由两位标注者调和标签的五篇论文试点中,它在完成的33项评估中,有25项分配的风险水平高于标签。在标签到达之前固定的规则消除了大部分样本内差异;加权一致性仍然较低。ARGUS提供证据关联的风险报告,定位潜在弱点以供专家审查,而不裁决因果主张。代码和数据:此https URL
英文摘要
Difference-in-differences (DID) studies are widely used to evaluate climate policy, but assessing the evidence supporting their identification assumptions remains challenging. We introduce ARGUS, a structured language-model pipeline that audits reported evidence against an eleven-dimension assumption-implication-evidence rubric and abstains when relevant evidence cannot be retrieved. We evaluate ARGUS using injected flaws, economics papers, and a small pilot with reconciled labels. On the 11-flaw benchmark, ARGUS detects 73% of planted flaws, compared with 18% for a keyword-based pipeline. Across 26 economics papers, ARGUS abstains on about 40% of paper-dimension assessments for lack of retrievable evidence. In a five-paper pilot with labels reconciled by two annotators, it assigns a higher risk level than the labels on 25 of the 33 assessments it completes. A rule fixed before the labels arrived removes most of this in-sample; weighted agreement stays low. ARGUS provides evidence-linked risk reports that localize potential weaknesses for expert review, without adjudicating causal claims. Code and data: https://github.com/yonghongzhang-io/ARGUS
CommentsAccepted at ClimateNLP 2026, the 3rd Workshop on Natural Language Processing meets Climate Change (EMNLP 2026). 9 pages plus appendix (21 pages total), 6 figures, 15 tables