揭示控制
Revelation Control
浏览论文内容
中文总结 AI 辅助
该研究针对学习系统提出揭示控制理论,定义决策充分揭示等概念,给出代价调整分解准则,通过Qwen2.5-7B和Mistral-7B-v0.3实验验证相关结论,支持结构迁移而非数值迁移。
中文摘要 AI 辅助
揭示控制是这样一类问题:选择带代价的干预措施,仅在被揭示的状态差异能够改变重要决策时才揭示隐藏状态,同时单独考虑干预措施本身带来的任何有用进展。我们为学习系统发展这一理论,其中在已声明的当前信息下等价的状态会对未来训练产生不同响应,并倾向于不同的动作。该框架定义了决策充分揭示与揭示深度,将纯信息价值与生产性复用分离,将静态贝叶斯精炼嵌入依赖状态的延续价值,并给出了精确的代价调整分解准则:仅当共享标量摘要的状态处于带代价的停止/继续边界的两侧时,新增的浅层坐标才是决策非冗余的。我们还给出了与目标无关的模型专属实例化协议,并证明在无限制严重程度下,仅有限的停止翻转风险无法证明正期望效用。在Qwen2.5-7B和Mistral-7B-v0.3上,更深的未来学习探测具有正决策价值,且生产性复用产生了严格的等计算量效用优势。Qwen还为决策非冗余浅层可揭示性 regime 提供了证据;在Mistral中,仅在独立开发面板上拟合的标量延续架构,在不相交的目标面板上保留了正的家族wise调整下界,与所测试的架构族和分辨率内的标量决策充分性一致。该证据支持的是结构而非数值迁移:决策理论、代价核算、延续逻辑和评估协议可迁移,而经验代理、系数、阈值甚至所需的浅层状态维度可能是系统专属的。
英文摘要
Revelation Control is the problem of choosing priced interventions that reveal hidden state only insofar as the revealed distinctions can change a consequential decision, while accounting separately for any useful progress created by the intervention itself. We develop this theory for learning systems, where states equivalent under declared current information can respond differently to future training and favor different actions. The framework defines decision-sufficient revelation and revelation depth, separates pure information value from productive reuse, embeds static Bayes refinement into state-dependent continuation value, and gives an exact cost-adjusted factorization criterion: an additional shallow coordinate is decision-nonredundant only when states sharing a scalar summary lie on opposite sides of the priced Stop/Continue boundary. We also give a target-independent protocol for model-specific instantiation and prove that bounded stop-flip risk alone cannot certify positive expected utility under unrestricted severity. Across Qwen2.5-7B and Mistral-7B-v0.3, deeper future-learning probes have positive decision value and productive reuse yields strict equal-compute utility advantages. Qwen additionally provides evidence for a decision-nonredundant shallow revealability regime; in Mistral, a scalar continuation architecture fit only on an independent development panel retains positive familywise-adjusted lower bounds on a disjoint target panel, consistent with scalar decision sufficiency within the tested architecture family and resolution. The evidence supports structural rather than numerical transfer: the decision theory, cost accounting, continuation logic, and evaluation protocol transport, while empirical proxies, coefficients, thresholds, and even the required shallow state dimension may be system-specific.