AI 中文总结
本文将探索与停止的动态控制问题简化为带信息预算约束的静态凸规划,结合时间偏好曲率得出最优探索形态,并将框架应用于实物期权等三类场景。
AI 中文摘要
我们研究一类决策者:其在行动前会进行探索,即动态选择学习内容,而后停止探索并采取行动。我们首先将该动态控制问题简化为一个静态问题:任何探索与停止策略都等价于在每个时期满足一个信息预算约束的前提下,对停止状态与停止时间的联合分布进行选择,我们精确刻画了哪些分布是可实现的。简化后的问题是一个带有线性目标的凸规划;其对偶变量随时间对信息定价,最优策略会凹化扣除这些影子价格后的停止收益。决策者的时间偏好曲率决定了最优探索的形态:凸时间偏好会引发泊松探索,凹时间偏好会将停止限制在一个窗口内,窗口长度由延迟边际成本的离散程度控制——当窗口较短时,会强制要求初始阶段进行纯探索——而线性情况则处于这两种情况的边界。我们将该框架应用于实物期权、信息获取中的速度-精度权衡以及连续时间探索竞赛。
英文摘要
We study a decision-maker who explores --- dynamically choosing what to learn --- before stopping to act. We first reduce this dynamic control problem to a static one: any exploration-and-stopping strategy is equivalent to a choice of the joint distribution of the stopped state and the stopping time, subject to one information-budget constraint at each date, and we characterize exactly which distributions are attainable. The reduced problem is a convex program with a linear objective; its dual prices information over time, and the optimal policy concavifies the stopping payoff net of these shadow prices. The curvature of the decision-maker's time preference then governs the shape of optimal exploration: convex time preference induces Poisson exploration, concave time preference confines stopping to a window whose length is controlled by the dispersion of the marginal cost of delay --- forcing an initial phase of pure exploration when the window is short --- and the linear case lies at the boundary between them. We apply the framework to real options, to the speed--accuracy tradeoff in information acquisition, and to a continuous-time exploration contest.