arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

离散非线性系统中基于熵的结构识别的有限样本极限

Finite-Sample Limits of Entropy-Based Structure Identification in Discretized Nonlinear Systems

Pratishtha Shukla, James Nutaro

arXiv 2609.03074首次发表:更新:

发表机构

Oak Ridge National Laboratory(橡树岭国家实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对离散非线性系统中基于熵的结构识别问题,引入分辨率-随机性比率,明确其有限样本极限,验证了理论并在配电网数据集上演示其用于识别驱动可靠性提升的投资。

AI 中文摘要

离散化从根本上限制了随机系统中的结构识别。当系统随机性超过离散化分辨率时,基于熵的方法将失去区分哪个输入驱动输出的能力。我们在模糊归纳推理(FIR,一种从离散测量中学习动力系统的非参数框架)中研究这一问题,其中输入变量的选择既决定预测精度,也决定所学输入-输出关系的可解释性。基于熵的选择针对可解释性,即识别哪些变量因果驱动输出;而基于均方误差(MSE)的选择针对预测。我们引入一个分辨率-随机性比率,该比率决定基于熵的选择何时可靠。由此得到三个结果:第一,在该阈值以下,基于熵的选择是一致的;但在阈值以上,无论样本量多大,它都会失去判别能力。第二,将熵选择的变量用于预测而非MSE选择的变量,会产生一个闭式形式的额外预测风险,该风险随输入复杂性增长,随样本量缩小。第三,可靠识别因果相关输入所需的数据量与输入组合的数量成正比,与熵信号的强度成反比。该理论在两状态马尔可夫模型上得到验证,并在分析基础设施投资影响的配电网可靠性数据集上得到演示,该研究的目标是解释哪些投资驱动可靠性提升,而非仅预测结果。

英文摘要

Discretization fundamentally limits structure identification in stochastic systems. When system stochasticity exceeds the discretization resolution, entropy-based methods lose their ability to distinguish which input drives the output. We study this in Fuzzy Inductive Reasoning (FIR), a nonparametric framework for learning dynamical systems from discretized measurements, where the choice of input variables determines both predictive accuracy and the interpretability of the learned input--output relationships. Entropy-based selection targets explainability, i.e., identifying which variables causally drive the output, while mean-squared-error-based selection targets prediction. We introduce a resolution-stochasticity ratio that governs when entropy-based selection is reliable. Three results follow. First, entropy-based selection is consistent below this threshold but loses discriminative power above it, regardless of sample size. Second, using the entropy-selected variables for prediction instead of the MSE-selected ones incurs a closed-form excess prediction risk that grows with input complexity and shrinks with sample size. Third, reliable identification of the causally relevant inputs requires data that scales with the number of input combinations and inversely with the strength of the entropy signal. The theory is validated on a two-state Markov model and demonstrated on a distribution grid reliability dataset analyzing the impact of infrastructure investment, where the goal is to explain which investments drive reliability improvements rather than merely predict outcomes.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑