arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.31197cs.LG

ALF:用于科学发现的活动学习框架

ALF: An Active Learning Framework for Scientific Discovery

  • InstaDeep
  • University College London(伦敦大学学院)
  • University of Oxford(牛津大学)

机构由 AI 辅助整理,请以论文原文为准。

Shikha Surana, Alex Hawkins-Hooker, Olivia Gallup, Christoph Brunken, Jules Tilly, Paul Duckworth

中文总结 AI 辅助

ALF是一个模块化的活动学习框架,通过五个组件统一离线基准测试与在线部署,在预算约束下高效获取高质量数据,推动科学发现。

中文摘要 AI 辅助

机器学习在科学发现中的应用几乎系统性地受限于数据。在预算约束下生成相关的高质量数据,是推动该领域发展的最有前景的途径之一。在标注需要昂贵实验、测量或模拟的场景中,活动学习(AL)提供了解决方案。现有的大多数工具仅覆盖数据获取循环的一部分,且通常侧重于离线基准测试或在线部署,而非两者兼顾。我们提出了ALF,一个模块化的活动学习框架,通过五个模块化组件运行完整的数据获取循环。它提供了一套统一的API,适用于两种场景:离线场景下,针对现有数据集进行受控且可复现的实验;在线场景下,针对真实部署中的预言机获取新候选。ALF是开源的,可在该https URL获取。

英文摘要

Machine learning for scientific discovery is almost systematically data bound. Producing relevant high quality data, under budget constraints, is amongst the most promising ways to advance the field. Active learning (AL) offers promise wherever labelling requires expensive experiment, measurement, or simulation. Most existing tools cover only part of the data acquisition loop, and typically focus on either offline benchmarking or online deployment, but not both. We present ALF, a modular AL Framework that runs the full data acquisition loop via five modular components. One clear API for both settings: offline, against an existing dataset for controlled and reproducible experimentation; and online, against an oracle for acquiring new candidates in real-world deployments. ALF is open-source and available at https://github.com/instadeepai/alf.

补充信息

↑