arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

IdeaAMBIG:研究思想规范中实现关键差距的基准测试

IdeaAMBIG: Benchmarking Implementation-Critical Gaps in Research-Idea Specifications

Yiling Ma, Yilun Zhao, Sihong Wu, Manasi Patwardhan, Arman Cohan

arXiv 2609.10539首次发表:更新:

发表机构

Yale University; TCS Research(耶鲁大学; 塔塔咨询服务研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

IdeaAMBIG基准测试660个实例,评估研究规范编码就绪性,发现缺陷定位是主要瓶颈,提供缺陷可大幅提升澄清成功率。

AI 中文摘要

一个研究思想可能具有新颖性、连贯性和科学合理性,但其提出的方法可能仍不足以被忠实地实现。我们研究了面向实现的研究方法规范的编码就绪性,该就绪性定义为规范是否为有能力的实现者或编码代理提供了足够的方法论信息,使其无需无依据的假设即可构建预期方法。我们从论文、代码库、问题线程和复现工件中构建了基于证据的规范及其支持的解决方案。我们引入了IdeaAMBIG,一个包含660个基于证据实例的基准:163个来自可复现性报告和GitHub问题的真实世界差距,以及497个注入到编码就绪参考中的受控合成差距。IdeaAMBIG评估三种能力:编码就绪性评估、缺陷定位和澄清行动生成。缺陷定位仅接收规范,而澄清额外接收标注的缺陷。在13个LLM中,最佳模型在真实世界实例上达到9.6%的宏缺陷恢复率,但在给定缺陷时达到80.6%的宏澄清行动成功率。在预言机研究中,提供黄金解决方案将下游编码就绪率从14%提高到98%。在所有评估模型中,缺陷定位是主要瓶颈,而在给定缺陷时澄清能力更强。

英文摘要

A research idea may be novel, coherent, and scientifically plausible, yet its proposed method may remain insufficiently specified for faithful implementation. We study the codification readiness of implementation-facing research-method specifications, defined by whether they provide sufficient methodological information for a competent implementer or coding agent to construct the intended method without unsupported assumptions. We construct evidence-grounded specifications and their supported resolutions from papers, codebases, issue threads, and reproduction artifacts. We introduce IdeaAMBIG, a benchmark of 660 evidence-grounded instances: 163 real-world gaps from reproducibility reports and GitHub issues, and 497 controlled synthetic gaps injected into codification-ready references. IdeaAMBIG evaluates three capabilities: codification-readiness assessment, defect localization, and clarification action generation. Defect localization receives only the specification, whereas clarification additionally receives the annotated defect. Across 13 LLMs, the best model achieves 9.6% Macro Defect Recovery Rate on real-world instances but 80.6% Macro Clarification Action Success Rate when given the defect. In an oracle study, supplying the gold resolution raises the downstream codification-ready rate from 14% to 98%. Across all evaluated models, defect localization is the main bottleneck, with stronger clarification given the defect.

CommentsPreprint. 74 pages, 18 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑