arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.07206cs.LGmath.OC

几何-非几何优化器演算:一种用于可达梯度方法的模块化语言

Causal Optimizer Interaction Calculus: Hidden Geometric Relaxation and Identifiable Interventions

Zavier Li

首次发表
浏览论文内容

中文总结 AI 辅助

研究提出几何-非几何优化器演算,用于审查可达梯度方法。核心方法是定义模块及相关机制,证明方向表达定理等。主要贡献是给出方向级诊断,将设计问题转为帕累托优化,还提升残差复杂性并提供诊断原型证据。

中文摘要 AI 辅助

自适应优化器混合了多种机制:度量或预处理器将梯度映射到下降方向,而估计、记忆、步长控制、约束、随机性、目标修改和离散化决定了哪些方向可用以及如何使用它们。我们引入了几何-非几何优化器演算,这是一种用于在显式预言机、预算、状态和规则约束下审查可达梯度方法的模块化语言。几何模块是一个正余度量族,它将余向量映射到参数空间方向;非几何模块是信息、记忆、控制、算子、噪声、目标和离散化机制。主要形式结果是一个方向表达定理:远离临界点时,完全正定几何精确地表达了严格下降方向。然后,我们为可允许的度量族定义了受限方向残差,证明了对角和块几何的精确表达条件,并将这种方向级诊断与条件数几何复杂性分开。由此产生的设计问题是对模块预算的帕累托优化,而不是单个通用优化器排序。我们还将逐点残差提升到轨迹级残差复杂性,将方向不匹配与解释几何的变化耦合起来。我们仅将诊断原型作为该语言的证据:一个高信息全度量探测器以数值精度解决确定性二次基准问题,而一个实用的μ子风格的PyTorch候选者提供了小规模证据,表明矩阵算子更新可以通过演算进行审查。本文是一篇理论和基准语言手稿;它并不声称具有大规模优化器的最新性能。

英文摘要

Optimizer experiments observe responses to algorithmic configurations without uniquely revealing hidden mechanisms. We develop a causal optimizer interaction calculus that separates pathwise realization, Mobius decomposition, and experimental identification. Under a fixed innovation coupling, every finite-horizon innovation-driven optimizer admits a behaviorally minimal pathwise realization. For any finite effect support and intervention design, an incidence operator gives the complete observational gauge, exact identifiability, sharp quotient stability, held-out predictions, and exact noiseless configuration complexity. Smooth hidden relaxation generates interactions through inverse hidden-state stiffness. Building on this structural law, we prove an observable-readout transfer theorem: arbitrary smooth update or trace readouts inherit an explicit five-term interaction through first and second hidden responses. Unlike the reduced optimal value, a general readout has no universal interaction sign. Its Boolean effects remain exact integrals of continuous interaction curvature and can therefore be identified by factorial interventions. We also derive Gaussian quotient minimax risk, exact confidence sets and tests, misspecification decomposition, certified downstream decisions, and optimal replication. A controlled real-data experiment on a 65-dimensional strongly convex logistic model validates the complete reduced-value chain. Boolean effects and independently integrated curvature agree within 4.21e-11, while nine held-out continuous intensities agree within 8.88e-13. Gaussian campaigns attain the predicted coverage and power, and 4,500 real-minibatch observations reject an order-two interaction model. Neural trace audits provide complementary evidence that the declared response classes remain informative in nonconvex training.

发表机构

  • Xidian University(西安电子科技大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑