AI 中文总结
该研究将算法决策形式化为终端计算分配问题,通过贝尔曼方程刻画最优计算分配,关联计算价值与信息,发现最大化近似VOC可得到加权A*,明确了无普遍最优获取规则的共享决策问题。
AI 中文摘要
许多算法在返回决策前会消耗内部资源,且仅通过最终输出的质量进行评估。我们将此类过程形式化为终端计算分配问题:高成本计算生成观测结果、更新对潜在环境的信念,且仅通过终端决策损失产生影响。贝尔曼方程刻画了在固定预算、定价计算和精确认证下的最优分配。随后我们将计算价值(VOC)与信息关联:在对数损失下,互信息等于近视VOC;在简单遗憾下,VOC是知识梯度量;此外,信息增益对计算的排序效果可能极差,但它给出了VOC的单侧上界。老虎机拉杆、树模拟和节点扩展在不同计算拓扑下阐释了同一模型。最后,在明确的前沿分辨率和启发式误差模型下,最大化近似VOC可得到加权A*,其中A*和贪心最佳优先搜索为极限情况。该理论确定了一个共享决策问题,而非断言某一获取规则普遍最优。
英文摘要
Many algorithms spend an internal resource before returning a decision and are evaluated only by the quality of that terminal output. We formalize such procedures as terminal computation-allocation problems: costly computations produce observations, update beliefs about a latent environment, and matter only through terminal decision loss. Bellman equations characterize optimal allocation under fixed budgets, priced computation, and exact certification. We then relate value of computation (VOC) to information. Mutual information equals myopic VOC under log loss, whereas under simple regret VOC is a knowledge-gradient quantity; moreover, information gain can rank computations arbitrarily poorly, although it gives a one-sided upper bound on VOC. Bandit pulls, tree simulations, and node expansions illustrate the same model under different computation topologies. Finally, under an explicit frontier-resolution and heuristic-error model, maximizing approximate VOC recovers weighted A*, with A* and greedy best-first search as limiting cases. The theory identifies a shared decision problem without asserting that one acquisition rule is universally optimal.