发表机构
Fondazione Bruno Kessler; IRVAPP(布鲁诺·凯塞勒基金会; IRVAPP)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对社会科学因果研究设计的人工评估依赖问题,本研究提出ARDTrA任务,构建专家标注数据集,用多轮RAG对话管道评估,发现段落长度是性能主因,且人机任务难度来源独立。
AI 中文摘要
社会科学中因果研究设计的可靠评估对循证决策至关重要,但目前完全依赖人工专家分析。我们提出了自动研究设计追踪与评估(ARDTrA)任务,该任务涉及检测论文中使用的研究设计并评估其应用质量。我们创建了由专家标注的论文数据集,涵盖六类反事实研究设计,并使用基于多轮检索增强生成(RAG)的对话式管道对该任务进行评估。在四种检索策略、四种大型语言模型(LLM)和六种嵌入模型的测试中,我们发现段落长度是性能的主要驱动因素,可解释52%-66%的方差。针对每种研究设计的分析还表明,人类与机器的难度并不一致:系统最难处理的设计并非专家标注者分歧最大的设计,这指向任务难度的两个独立来源。
英文摘要
Reliable assessment of causal research designs in the social sciences is critical for evidence-based policy-making, yet has so far relied entirely on manual expert analysis. We introduce Automated Research Design Tracking and Assessment (ARDTrA), a task that involves detecting the research design used in a paper and assessing the quality of its application. We create an expert-annotated dataset of papers covering six families of counterfactual research designs and evaluate the task using a multi-turn RAG-based conversational pipeline. Across four retrieval strategies, four LLMs and six embedding models, we find that passage length is the main driver of performance, explaining 52-66% of the variance. A per-research-design analysis also shows that human and machine difficulty do not align: the designs that prove hardest for the system are not those on which expert annotators disagree most, pointing to two independent sources of task difficulty.
CommentsPaper accepted at EMNLP 2026 - Main Conference