结构而非信念:组合半臂老虎机中基于LLM导出协方差的关联汤普森采样
Structure, Not Belief: Correlated Thompson Sampling from LLM-Derived Covariance in Combinatorial Semi-Bandits
浏览论文内容
中文总结 AI 辅助
针对组合半臂老虎机,提出仅用LLM划分臂结构生成协方差矩阵来关联汤普森采样,不注入信念,理论证明并实验显示可显著降低遗憾。
中文摘要 AI 辅助
组合汤普森采样(CTS)对每个臂独立抽取后验样本,因此其探索动态忽略了臂之间的任何关联。我们研究了这些动态的一个最小改动:对臂的一个划分仅查询一次LLM,该划分通过聚类秩上的RBF核成为一个正定相关矩阵$\Sigma$,并且每轮后验样本以协方差$\Sigma$抽取,而Beta后验仅从真实奖励更新,因此LLM塑造的是采样器的移动方式,而非其信念。我们为理想化的高斯采样器给出了一个自包含的贝叶斯遗憾界,其信息增益分解为来自$K$聚类结构的$K\log T$项和一个增长至$d\log T$的岭项:相对于独立采样的$\sqrt{d/K}$改进是一个有限时域瞬态,仅在簇内相关性趋于1时精确成立。在$T=2{,}500$时,相关采样器在16个合成伯努利族上将遗憾比CTS降低了19%(在$T=25{,}000$时使用数据自适应核降低6-7%),在微软MIND-small新闻基准($d=200$篇真实文章)上降低了41%,而伪观测热启动无任何收益。一个无LLM的消融实验使用质量受控的模拟预言机表明,在非结构化实例上,增益是核形状的属性(随机划分或对采样噪声的简单温度调节即可复现),而在匹配预言机质量下注入信念从未有帮助。
英文摘要
Combinatorial Thompson sampling (CTS) draws independent posterior samples for every arm, so its exploration dynamics ignore any relation among arms. We study a minimal change to those dynamics: an LLM is queried once for a partition of the arms, the partition becomes a positive-definite correlation matrix $Σ$ through an RBF kernel on cluster ranks, and the per-round posterior sample is drawn with covariance $Σ$ while the Beta posteriors are updated from real rewards only, so the LLM shapes how the sampler moves, not what it believes. We give a self-contained Bayesian regret bound for the idealized Gaussian sampler whose information gain splits into a $K\log T$ term from the $K$-cluster structure and a ridge term that grows to $d\log T$: the $\sqrt{d/K}$ improvement over independent sampling is a finite-horizon transient, exact only as the within-cluster correlation tends to one. The correlated sampler reduces regret by 19% over CTS on 16 synthetic Bernoulli families at $T=2{,}500$ (6-7% at $T=25{,}000$ with data-adaptive kernels) and by 41% on the Microsoft MIND-small news benchmark ($d=200$ real articles), while pseudo-observation warm starts give nothing. An LLM-free ablation with a simulated oracle of controlled quality shows that on unstructured instances the gain is a property of the kernel shape (a random partition, or a plain tempering of the sampling noise, reproduces it), while belief injection at matched oracle quality never helps.
发表机构
- Princeton University(普林斯顿大学)
机构由 AI 辅助整理,请以论文原文为准。