一种辨别演算:决策相关洞察、序列价值与遗忘作为高阶学习
A Calculus of Discernment: Decision-Relevant Insight, Sequence Value, and Forgetting as Higher-Order Learning
- Independent AI Researcher, Zürich, Switzerland(独立AI研究员,瑞士苏黎世)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
研究在生成式AI丰富候选洞察下,稀缺的辨别、行动及遗忘能力问题。构建围绕特定对象的框架,提出APOHA理论。通过在肥胖治疗决策世界测试,表明自适应遗忘能降决策后悔、优化记忆,价值感知遗忘有优势。
AI中文摘要:
在生成式人工智能的世界中,候选洞察丰富,但辨别哪些重要、以正确的数量和顺序对其采取行动以及遗忘其余部分以使系统能够适应的能力却很稀缺。我们认为这些稀缺性受一个对象支配,并围绕它构建了一个框架。我们将洞察严格定义为对目标有确定、可衡量影响的杠杆,并通过信息的期望值而非新颖性按决策相关性对候选进行排序。我们表明行动不仅有大小,还有顺序:在现实的信念动态下,内容‘触碰’是非对易算子,所以以不同顺序交付的固定计划会产生不同结果,定义了序列溢价。我们观察到任何杠杆的价值都是影子价格,将药物营销、股票选择和制造统一为一个杠杆发现问题。最具推测性的是,我们提出了APOHA理论,其中遗忘不是知识的处置,而是学习价值的算子:保留项目的价值是遗忘它的反事实成本,学习系统是在保留价值的前提下最大遗忘的残余,高阶价值是在反复遗忘后幸存的结构(与重整化相关的不变量)及巩固作为其共轭。我们陈述了核心开放问题(具有谱隙的非平凡吸引子)并测试了遗忘理论:将APOHA作为一个智能体应用于非平稳肥胖治疗决策世界,在30个种子上,自适应遗忘与从不遗忘和固定半衰期相比,将累积决策后悔降低了24 - 32%,保持了约6倍小且更干净的记忆,并稳定收敛;值得注意的是,盲目遗忘比从不遗忘更糟,所以这种好处特定于价值感知遗忘。多学科批评对整体进行了压力测试。
英文摘要:
In a world of generative AI, candidate insights are abundant; what is scarce is the capacity to discern which matter, to act on them in the right amount and order, and to forget the rest so the system can adapt. We argue these scarcities are governed by one object and build a framework around it. We define an insight strictly as a lever with an identified, measurable effect on an objective, and rank candidates by decision-relevance via the expected value of information rather than novelty. We show action carries an order, not only a size: under realistic belief dynamics, content "touches" are non-commuting operators, so a fixed plan delivered in different orders yields different outcomes, defining a sequence premium. We observe that the value of any lever is a shadow price, unifying pharmaceutical marketing, equity selection, and manufacturing as one leverage-discovery problem. Most speculatively, we propose APOHA, a theory in which forgetting is not the disposal of knowledge but the operator by which value is learned: the value of a retained item is the counterfactual cost of forgetting it, a learning system is the residue of maximal forgetting subject to preserved value, and higher-order value is the structure that survives repeated forgetting (a renormalisation-relevant invariant), with consolidation as its conjugate. We state the central open problem (a non-trivial attractor with a spectral gap) and test the forgetting theory: operationalising APOHA as an agent on a non-stationary obesity-treatment decision world over 30 seeds, adaptive forgetting cut cumulative decision-regret by 24-32% against never-forget and a fixed half-life, kept a ~6x smaller, cleaner memory, and converged stably; notably, blind forgetting was worse than never forgetting, so the benefit is specific to value-aware forgetting. A multi-disciplinary critique stress-tests the whole.