发表机构
Georgetown University; Nokia Bell Labs(乔治城大学; 诺基亚贝尔实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究以电信工单检索为例,探讨自主研究在开放式工业级ML问题中的应用,发现其虽缺乏人类直觉,但能在低成本和短时间内达到90%的最优性能,建议人机协同研究。
AI 中文摘要
基于LLM的系统在问题求解和编码方面的近期突破,推动了AI for Science范式的发展,有望取代机器学习(ML)研究中的人类角色。然而,尽管已提出多个全自主端到端ML研究框架,其成功实施往往局限于搜索空间狭窄的问题,如语言建模或生物医学ML基准。本文通过一个案例研究——电信工单检索,探讨如何将自主研究应用于解决开放式的、工业级ML问题。该任务在表示、架构和训练数据生成方面具有自由度。我们发现,使用商业和开源智能体的自主研究在解决开放式问题时既展现出潜力也存在局限:自主研究虽能擅长窄范围的超参数优化,但缺乏类人的直觉和创造力,且需要运营开销。即使在最少人工监督下,自主研究也能在更短时间内(10周对比人工10个月)以适度成本(每次Cursor活动最高200美元)达到最先进性能的90%(Recall@1为0.34对比0.38)。我们的实证证据建议,人类研究人员与自主研究框架协同工作,以实现ML研究的最佳效果。
英文摘要
Recent breakthroughs in LLM-based systems and their abilities in problem solving and coding have allowed progress in the AI for Science paradigm, potentially replacing human roles in machine learning (ML) research. However, while several frameworks of fully autonomous end-to-end ML research have been proposed, successful implementations of them are often limited to problems with narrow search spaces, like language modeling or biomedical ML benchmarks. In this paper, we explore how autonomous research can be adapted to solve open-ended, industry-grade ML problems, by considering a case study: telecom ticket retrieval, an open-ended task with degrees of freedom in representation, architecture, and training data generation. We discover that autonomous research for open-ended problems with commercial and open-source agents shows both promise and limitations: while autonomous research can excel in narrow hyperparameter optimization, it lacks human-like intuition and creativity and requires operational overhead. Even with minimal human supervision, autonomous research can reach $90\%$ of state-of-the-art performance (0.34 vs. 0.38 Recall@1) in a much shorter time period (10 weeks vs. 10 months of human work) at a modest cost (up to \$200 per Cursor campaign). Our empirical evidence recommends that human researchers and autonomous research frameworks work together for best results in ML research.
Comments10 pages, 3 tables