AI 中文总结
研究在纵向数据准备任务中评估开放权重的大语言模型,介绍开源框架,含真实数据集、任务定义和评估程序,通过基准测试发现当前先进模型表现良好,为治理受限研究中的AI辅助数据准备提供可行路径。
AI 中文摘要
大语言模型(LLMs)和智能体在代码开发中广泛应用,数据常发送到第三方云模型。在使用个人数据的研究中,治理要求限制了其采用。本地可部署的开放权重模型提供了替代方案。本文介绍一个开源框架,用于评估由开放权重LLMs驱动的人工智能智能体在纵向人口研究数据准备这一持续瓶颈任务上的效果。该框架包括精心策划的真实数据集、任务定义和评估LLM生成的R代码及输出数据的自动化程序。通过对LLMs在20个数据准备任务上进行基准测试,当前最先进的31 - 35B参数模型几乎达到基准上限。开放权重LLMs在消费级硬件上的性能为治理受限研究环境中的人工智能辅助数据准备提供了可行途径。
英文摘要
Large language models (LLMs) and agents are now widely used tools in code development, with data typically sent to third-party cloud-based models. Their adoption in research using personal data is constrained by governance requirements that typically prohibit data transmission to external services. Locally deployable open-weight models offer an alternative since sensitive data never leave the local environment. We introduce an open-source framework for evaluating the efficacy of AI agents powered by open-weight LLMs on one of the most persistent bottlenecks in research on longitudinal population studies: data preparation. The framework comprises: a curated ground-truth dataset (cleaning scripts preparing six sweeps of data from a British cohort study), task definitions encompassing tasks such as category harmonization and multi-wave merging, and automated routines for evaluating the LLM-produced R code and outputted data. We benchmark LLMs across the (consumer grade) deployment spectrum to assess their efficacy in 20 data preparation tasks (creation of 102 variables). Current state-of-the-art, 31-35B parameter models almost saturated our benchmark ('average task completion' up to 87.9%). The performance of open-weight LLMs running on consumer-grade hardware shows promise of a viable path toward AI-assisted data preparation in governance-restricted research settings. Our framework is publicly available at: https://github.com/UCL-ARC/RRBench.
CommentsPresented at SLLS 2026; accepted at CLS 2026 and RSS 2026