arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.13504cs.LG

稀疏正交回归技术:用于方程发现、近似与积分的谱框架

Sparse Orthogonal Regression Technique: A Spectral Framework for Equation Discovery, Approximation, and Integration

Sabin Roman, Ljupco Todorovski, Saso Dzeroski

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出SORT稀疏正交回归技术,作为从含噪不规则采样数据学习正交基展开的稀疏谱框架,可用于方程发现、近似与积分,在动力系统实验中表现优于基线方法,性能下降更稳定。

中文摘要 AI 辅助

我们开发了稀疏正交回归技术(Sparse Orthogonal Regression Technique,SORT),这是一种从含噪且不规则采样的数据中学习正交基展开的稀疏谱框架。SORT使用L1正则化回归直接从观测值估计展开系数,避免了显式求积或解析内积评估。其核心应用是数据驱动的常微分方程发现:向量场以选定的正交基表示,并作为稀疏系数展开学习。这为符号回归、基于语法的发现以及SINDy式稀疏识别提供了一条互补途径,即先恢复紧凑的谱表示,后续可指导更简单解析形式的搜索。在动力系统实验中,当基函数适配问题时,SORT与基于库的稀疏回归基线方法表现相当或更优;在稀疏采样、含噪导数估计及表示不匹配的情况下,其性能下降更稳定。具体示例阐明了该表示的效用:若有限库缺失问题特定的非线性,所得模型可能失效。SORT并非对不匹配免疫,但它将问题从通用项间脆弱选择转向适配问题域的基函数设计。实验还表明,主导低阶系数随模型阶数增加而持续存在,支持阶数一致的模型增长。除方程发现外,相同的学习展开还可通过系数读出支持非线性近似及复杂高维积分的估计。总体而言,SORT为系统识别、近似与积分提供了可复用的中间表示,同时将基函数设计明确化为科学建模问题的一部分。

英文摘要

We develop the Sparse Orthogonal Regression Technique (SORT), a sparse spectral framework for learning orthonormal-basis expansions from noisy and irregularly sampled data. SORT estimates expansion coefficients directly from observations using L1-regularized regression, avoiding explicit quadrature or analytic inner-product evaluation. The central application is data-driven discovery of ordinary differential equations: vector fields are represented in chosen orthogonal bases and learned as sparse coefficient expansions. This provides a complementary route to symbolic regression, grammar-based discovery, and SINDy-style sparse identification by first recovering a compact spectral representation, which can later guide searches for simpler analytic forms. Across the dynamical-system experiments, SORT matches or improves upon library-based sparse-regression baselines when the basis is well adapted to the problem, and shows more stable degradation under sparse sampling, noisy derivative estimates, and representation mismatch. Specific examples illustrate why this representation is useful: if a finite library misses the problem-specific nonlinearity, the resulting model can fail. SORT is not immune to mismatch, but it shifts the problem away from brittle selection among generic terms to basis design adapted to the problem domain. The experiments also show that dominant low-order coefficients persist as model order increases, supporting order-consistent model growth. Beyond equation discovery, the same learned expansion supports nonlinear approximation and estimation of complex, high-dimensional integrals by coefficient readout. Overall, SORT provides a reusable intermediate representation for system identification, approximation, and integration, while making basis design an explicit part of the scientific modeling problem.

发表机构

  • Jožef Stefan Institute(约热夫·斯泰凡研究所)
  • University of Ljubljana(卢布尔雅那大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑