arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

该选哪个LLM?面向大语言模型的在线主动模型选择

Which LLM to pick? Online Active Model Selection for Large Language Models

Alessandro Turrin, Patrik Okanovic, Torsten Hoefler, Nezihe Merve Gürel

arXiv 2610.01592首次发表:更新:

发表机构

TU Delft; ETH Zurich(代尔夫特理工大学; 苏黎世联邦理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对流式数据下LLM选型依赖基准不准确且人工标注昂贵的问题,提出ONLINE LLM PICKER框架,以有限标注预算主动选择信息量最大的提示,在130余模型上节省最高71.67%标注成本并降低2.51倍遗憾值。

AI 中文摘要

大语言模型(LLMs)越来越多地被应用于处理流式数据,从业者依赖基准测试来选择最佳模型,尽管这些信号只能近似真实性能。虽然人工标注可以提供可靠的反馈,但通常成本高昂且难以大规模获取。为应对这一挑战,我们提出了ONLINE LLM PICKER,这是首个在在线场景下针对LLMs进行主动模型选择的框架。给定任意查询流和有限的标注预算,ONLINE LLM PICKER选择最具信息量的提示进行标注,以在候选模型中识别出最佳LLM。在包括10个数据集的多个任务中,针对超过130个语言模型,我们展示了ONLINE LLM PICKER最多可节省71.67%的标注成本,同时可靠地识别出该数据流中最佳或接近最佳的模型。我们还表明,使用返回的模型对未标注提示进行顺序生成,可将遗憾值最多降低2.51倍,这表明ONLINE LLM PICKER能够在处理完所有流式提示之前就识别出最佳或接近最佳的模型。

英文摘要

Large Language Models (LLMs) are increasingly applied to process streaming data, with practitioners relying on benchmarks to select the best model even though these signals only approximate real performance. While oracle annotations can provide reliable feedback, they are often costly and difficult to obtain at scale. To address this challenge, we propose ONLINE LLM PICKER, the first framework for active model selection for LLMs in online settings. Given an arbitrary stream of queries and a limited annotation budget, ONLINE LLM PICKER selects the most informative prompts for annotation to identify the best LLM among candidate models. Across multiple tasks including 10 datasets, for over 130 language models, we show that ONLINE LLM PICKER saves annotation cost by up to 71.67% while reliably identifying the best or near-best model for the stream. We also show that using the returned model for sequential generation on unannotated prompts across the stream reduces regret by up to a factor of 2.51x, indicating that ONLINE LLM PICKER can identify the best or near-best model well before processing all streaming prompts.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑