arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

生产环境中的多臂老虎机:推理时的超参数优化

Bandits in Prod: Hyperparameter Optimization at Inference Time

Louis Abraham, Tuan-Anh Nguyen, Nicolas Devatine

arXiv 2609.01335首次发表:更新:

发表机构

Tiime(Tiime)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对生产系统推理时超参数优化的问题,提出IMABO框架,结合IMOSS策略与三类神谕,在多类OHPO场景中实现了更低的累积遗憾。

AI 中文摘要

许多生产系统只能通过在实时请求中使用配置并观察有噪声的反馈来评估该配置。现代智能体系统就是一个典型例子,其推理时的选择包括模型选择、检索深度、提示策略和解码温度,但往往没有代表性的验证数据。我们将此场景形式化为在线超参数优化(OHPO),并将其转化为混合条件搜索空间上的无穷多臂老虎机问题。我们提出IMABO,这是一个通用框架,它结合了任意用于在已采样配置中进行选择的老虎机策略,以及任意用于提出新配置的神谕。我们用IMOSS实例化该框架,这是一种无重启的任意时间策略,其活跃集随$t^\beta$增长,其中$\beta\in(0,1)$控制活跃集的增长,$p_\rho$是提议配置落在搜索空间前$\rho$分位的概率下界。我们将IMOSS与三个实用神谕结合:树状Parzen估计器、由坐标级老虎机驱动的当前最优变异神谕,以及预训练表格基础模型,这三者均优于均匀随机神谕基线。IMABO在从调优经典机器学习模型到配置基于LLM的智能体的各类OHPO场景中,获得了最低的累积遗憾。

英文摘要

Many production systems can assess a configuration only by using it on live requests and observing noisy feedback. Modern agentic systems are a prominent example, with inference-time choices such as model selection, retrieval depth, prompting strategy, and decoding temperature, yet often with no representative validation data. We formalize this setting as Online Hyperparameter Optimization (OHPO) and cast it as an infinitely many-armed bandit over mixed and conditional search spaces. We introduce IMABO, a general framework that combines any bandit policy for choosing among already sampled configurations with any oracle for proposing new ones. We instantiate it with IMOSS, a restart-free anytime policy whose active set grows as $t^β$, and prove an expected cumulative quantile-regret bound of $O(p_ρ^{-1/β} + T^{(1+β)/2})$, where $β\in(0,1)$ controls active-set growth and $p_ρ$ lower-bounds the probability that a proposed configuration falls in the top-$ρ$ fraction of the search space. We combine IMOSS with three practical oracles: a Tree-structured Parzen Estimator, an incumbent-mutation oracle driven by a per-coordinate bandit, and a pretrained tabular foundation model, all three improving over the uniform random oracle baseline. IMABO outperforms all baselines in terms of regret across diverse OHPO settings, from tuning classical machine-learning models to configuring LLM-based agents. Our implementation is available at https://github.com/Tiime-Software/IMABO.

Comments32 pages, 14 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑