InterveneBench: Benchmarking LLMs for Intervention Reasoning and Causal Study Design in Real Social Systems
InterveneBench:用于干预推理和真实社会系统因果研究设计的LLM基准测试
机构 * Fudan University(复旦大学) ; Shanghai Innovation Institute(上海创新研究院)
专题命中 推理评测 :reasoning(title,abstract);分类 cs.AI
AI总结 InterveneBench通过真实社会科学研究评估LLM在干预推理和因果研究设计中的能力,发现现有模型表现不足,并提出STRIDES多智能体框架提升性能。
Comments 35pages,3 figures