发表机构
Hanoi University of Science and Technology; Quantum AI & Cyber Security Institute, FPT Corporation(河内科技大学; FPT公司量子人工智能与网络安全研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出 VarDE 方法,将其应用于 BAI、MCTS、BPI 问题,在方差衰减和简单遗憾上有理论保证,在高度随机环境中性能优于现有方法。
AI 中文摘要
我们提出方差驱动探索(VarDE),这是一种用于高度随机环境中纯探索的原则性方法,此类环境下探索过程由随机方差主导。VarDE 基于一项基本原则:应分配采样精力以最小化最终决策的不确定性。我们通过平滑决策函数将最终决策的不确定性形式化,并推导分配规则,明确捕捉各组成部分的随机噪声如何影响最终输出的可靠性。我们将该方法应用于纯探索的三个核心问题——最佳臂识别(BAI)、蒙特卡洛树搜索(MCTS)和最优策略识别(BPI),并在方差衰减和简单遗憾方面提供理论保证。实验表明,与现有方法相比,VarDE 始终且显著地提升了性能,尤其在高度随机环境中表现出强劲增益。
英文摘要
We propose Variance Driven Exploration (VarDE), a principled approach for pure exploration in highly stochastic environments, where the exploration process is dominated by stochastic variance. VarDE is built on a fundamental principle: sampling effort should be allocated to minimize the uncertainty of the final decision. We formalize the uncertainty of the final decision through a smooth decision function and derive allocation rules that explicitly capture how stochastic noise in individual components affects the reliability of the final output. We apply this methodology to three core problems of pure exploration -- Best Arm Identification (BAI), Monte Carlo Tree Search (MCTS), and Best-Policy Identification (BPI) -- with theoretical guarantees on variance decay and simple regret. Empirically, we demonstrate consistent and significant improvements of VarDE over existing methods, with especially strong gains in highly stochastic environments.
CommentsTo appear in Proceedings of the 43rd International Conference on Machine Learning (ICML 2026)