arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.30017cs.LGcs.AI

Canopy:利用分段平滑树先验的多保真度老虎机

Canopy: Exploiting Piecewise Smooth Tree Priors for Multi-Fidelity Bandits

  • Amazon(亚马逊)
  • Columbia University(哥伦比亚大学)

机构由 AI 辅助整理,请以论文原文为准。

Michael Jerge, Suman Jana

AI总结:

针对多保真度树优化问题,提出CANOPY方法,通过在线学习分段平滑先验的有效区域,引导昂贵评估,在多种推理任务中显著提升性能。

AI中文摘要:

许多大语言模型推理问题,包括模型路由、前缀缓存管理、提示词裁剪和测试时搜索,都可以视为在树上进行优化。这种结构自然产生于自回归生成:每个前缀定义一个节点,其后续内容构成其下方的子树。树的内部节点提供对区域价值的廉价但有偏估计,而叶节点评估则昂贵但准确。分层老虎机方法可以利用这种结构,但通常需要预先指定特定的平滑度调度,即使真实目标往往只是分段平滑的,且其最优值可能位于尖锐边界附近。我们提出了CANOPY,一种多保真度树老虎机,它学习平滑先验在何处有效,而非全局假设其有效。CANOPY使用廉价的随机路径探测来构建局部聚合偏差的在线证书,然后将昂贵的叶节点评估导向证书检测到平滑度违规的单元。我们证明了固定预算和遗憾保证,其额外成本在间断点数量上是加性的,当不存在违规时恢复平滑树速率,当违规变得密集时接近无结构搜索。在路由、top-k识别、测试时搜索、缓存和提示词裁剪中,CANOPY持续改进匹配预算的性能,包括在1000个模型池上top-10召回率提高2.9倍,比best-of-N多解决1.6倍的SWE-bench Verified问题,以及使用前缀缓存时中位首令牌时间降低3.6倍。

英文摘要:

Many LLM inference problems, including model routing, prefix-cache management, prompt trimming, and test-time search, can be viewed as optimization over a tree. This structure arises naturally from autoregressive generation: every prefix defines a node, and its continuations form a subtree below it. Internal nodes of the tree provide cheap but biased estimates of a region's value, while leaf evaluations are expensive but accurate. Hierarchical bandit methods can exploit this structure, but typically require a specific smoothness schedule to be specified in advance, even though real objectives are often only piecewise smooth and their optima may lie near sharp boundaries. We introduce CANOPY, a multi-fidelity tree bandit that learns where the smoothness prior is valid rather than assuming it globally. CANOPY uses cheap random-path probes to construct an online certificate of local aggregation bias, then directs expensive leaf evaluations toward cells where the certificate detects a smoothness violation. We prove fixed-budget and regret guarantees whose additional cost is additive in the number of discontinuities, recovering the smooth-tree rate when no violations are present and approaching structure-blind search as violations become dense. Across routing, top-$k$ identification, test-time search, caching, and prompt trimming, CANOPY consistently improves matched-budget performance, including $2.9\times$ higher top-10 recall on a 1000-model pool, $1.6\times$ more SWE-bench Verified issues resolved than best-of-$N$, and $3.6\times$ lower median time-to-first-token with prefix caching.

↑