arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

成本感知的最优大语言模型识别:基于成对反馈

Cost-Aware Best-LLM Identification using Dueling Feedback

Sarvesh Gharat, Nikhil Karamchandani, Jayakrishnan Nair

arXiv 2609.30360首次发表:更新:

发表机构

Centre for Machine Intelligence and Data Science (C-MInDS); Indian Institute of Technology Bombay(机器智能与数据科学中心; 印度理工学院孟买分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对异构查询成本下的大语言模型最佳选择问题,提出基于成对反馈和成本感知的多臂老虎机算法,实现渐近最优成本并验证其有效性。

AI 中文摘要

受从一组具有异构查询成本的大语言模型(LLMs)中识别最佳模型这一问题的启发,我们提出并分析了一种多臂老虎机(MAB)的变体,该变体具有:(i) 成对反馈,其中模型响应之间的两两比较提供了稳健的偏好信号;(ii) 异构采样成本,反映了查询不同LLM的成本差异。在假设存在Condorcet赢家(我们在多个真实世界数据集上经验验证了这一条件)的前提下,我们提出了一种Track-and-Stop风格的算法,用于在给定置信度下进行最佳臂识别。我们证明了该算法在误差趋于零时几乎必然达到渐近最优成本。最后,我们在合成和真实世界实例上广泛评估了我们的方法,展示了相对于经典的成本无关算法及其成本感知扩展的一致改进。

英文摘要

Inspired by the problem of identifying the best model from a collection of large language models (LLMs) with heterogeneous querying costs, we formulate and analyse a variant of the multi-armed bandit (MAB) with (i) dueling feedback, where pairwise comparisons between model responses provide robust preference signals, and (ii) heterogeneous sampling costs, reflecting the differing costs of querying different LLMs. Assuming the existence of a Condorcet winner, a condition we empirically validate across multiple real-world datasets, we propose a Track-and-Stop style algorithm for best-arm identification with prescribed confidence. We prove that the algorithm almost surely achieves the asymptotically optimal cost as the error tends to zero. Finally, we extensively evaluate our approach on both synthetic and real-world instances, demonstrating consistent improvements over classical cost-unaware algorithms and their cost-aware extensions.

CommentsWe propose a cost-aware dueling bandit algorithm for best arm identification, prove its asymptotic optimality, and demonstrate its effectiveness in reliably identifying the best LLM with a minimum cost Accepted at NeurIPS 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑