AI 中文总结
研究学习环境中决策者面临的两难困境,提出RAEC算法平衡短期性能和长期承诺,建立遗憾上界及极小极大下界,还研究扩展及算法ROSCOC,通过数值实验验证算法优于基线算法。
AI 中文摘要
学习环境中的决策者面临两难困境,短期最优行动可能不利于长期利益。为理解此困境背后的基本权衡,研究后承诺奖励转移的自适应实验。实验阶段决策者可自适应测试多个选项,承诺阶段必须选定一个选项,其奖励可能与承诺前不同。提出保留臂消除承诺(RAEC)算法,预留部分实验阶段识别最佳转移后选项,其余轮次最小化短期遗憾。建立RAEC在所有参数范围的遗憾上界及匹配的极小极大下界。还研究了两个扩展,对于有先验结构知识的情况,正确识别转移的排名变化分量比估计其绝对大小更重要;对于具有凹承诺奖励和投资组合选择的设置,开发了保留在线随机凸优化承诺(ROSCOC)算法。最后进行数值实验,证实所提算法达到理论预测的遗憾,且优于其他基线算法。
英文摘要
Decision-makers in learning environments face a dilemma when their short-term optimal actions may not favor their long-term benefits the most. To understand the fundamental tradeoff behind the dilemma, we study adaptive experimentation with post-commitment reward shifts. During an experiment phase, the decision-maker may adaptively test multiple options; during a subsequent commitment phase, the decision-maker must commit to a single option, whose reward may differ from its pre-commitment reward. We propose the Reserved Arm Eliminations for Commitment (RAEC) algorithm, which reserves a predetermined portion of the experiment phase to identify the best post-shift option while using the remaining rounds to minimize short-run regret. We establish regret upper bounds for RAEC across all parameter regimes and matching minimax lower bounds, providing a tight characterization of the cost of balancing short-term performance and long-term commitment. A key implication is that deciding in advance how much of the experiment phase to reserve for the commitment decision is sufficient to achieve the best possible worst-case regret rate; adapting this amount as more data are observed does not improve the rate. We further study extensions with structural knowledge of reward shifts and with concave commitment rewards and portfolio choice. Numerical experiments confirm that our proposed algorithms achieve the regret predicted by our theory and outperform other baselines.