arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

能力门控规划:目标成本发现与近视实验选择的局限性

Capability-Gated Planning: Cost-to-Goal Discovery and the Limits of Myopic Experiment Selection

Ahmed Hassoon, Mark Dredze

arXiv 2608.05085首次发表:更新:

发表机构

Johns Hopkins University(约翰斯·霍普金斯大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究指出近视实验选择方法无法评估构建认知能力的价值,引入能力门控规划器 CG-Plan 解决此问题,证明近视规划器存在无界近似比且无法到达目标。

AI 中文摘要

自动化科学发现的系统必须反复决定要运行哪个实验、测试哪个假设、构建哪个工具以及何时停止。许多系统通过最大化近视分数来做出这些决策,例如每单位成本的预期信息增益或学习到的合理性分数。我们确定了这种方法的一个结构性局限性:一些行动具有建设性,它们获取一种认知能力(如仪器、检测方法、流程、模拟器或抽象),其价值不在于立即返回的信息,而在于它为未来行动提供的可能性。当通往可靠答案的最低成本路径需要一系列此类构建时,仅通过有限时间范围内可获得的信息来对行动评分的规划器无法对第一次构建进行估值,因为它在该时间范围内不会产生任何信息,且会被任何具有正信息的测量所主导,无论其多么微小。我们将目标导向的发现表述为信念空间中的随机最短路径问题,其中建设性实验会改变下游行动图,并证明对于每个展望深度 d,都存在一个实例,使得每个近视信息最大化规划器都具有无界近似比,且存在一个相关实例,使得它永远无法到达目标。该机制是一个能力不可区分性引理:在该时间范围内,获取能力在观察上可能与执行空操作无法区分。这确立了能力门控是一种与曲率(子模块性)和信息顺序(适应性差距)不同的可达性难度轴。我们引入了 CG-Plan,这是一种增量重规划器,具有感知能力的目标成本启发式 h = h_cap + h_exp。在受控测试平台中,性能差距仅在门控情况下出现,对每个固定展望都持续存在,且当近 miss 假设来自数据一致的提议者时会产生。

英文摘要

Systems that automate scientific discovery must repeatedly decide which experiment to run, which hypothesis to test, which tool to build, and when to stop. Many systems make these decisions by maximizing a myopic score such as expected information gain per unit cost or a learned plausibility score. We identify a structural limitation of this approach. Some actions are constructive: they acquire an epistemic capability (an instrument, assay, pipeline, simulator, or abstraction) whose value lies not in the information returned immediately but in the future actions it makes available. When the least-cost route to a confident answer requires a chain of such constructions, a planner that scores actions only by information obtainable within a bounded horizon cannot value the first construction: it yields no information within the horizon and is dominated by any measurement with positive information, however small. We formulate goal-directed discovery as a stochastic shortest-path problem in belief space in which constructive experiments change the downstream action graph, and prove that for every lookahead depth d there is an instance on which every myopic information-maximizing planner has an unbounded approximation ratio, and a related instance on which it never reaches the goal. The mechanism is a capability-indistinguishability lemma: within the horizon, acquiring a capability can be observationally indistinguishable from paying for a null action. This establishes capability gating as a reachability axis of difficulty distinct from curvature (submodularity) and information order (adaptivity gaps). We introduce CG-Plan, an incremental replanner with a capability-aware cost-to-go heuristic h = h_cap + h_exp. In a controlled testbed, the performance gap appears only under gating, persists for every fixed horizon, and arises when near-miss hypotheses come from a data-consistent proposer.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑