arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.09613stat.MEstat.ML

用于分布数据的带特征选择的动态弗雷歇回归

Dynamic Frechet Regression with Feature Selection for Distributional Data

Kiran Adhikari, Amrutha Dinesh, Mathew Kuttolamadom, Ying Lin

首次发表
浏览论文内容

中文总结 AI 辅助

针对分布数据回归难题,提出动态弗雷歇回归框架,通过索引感知加权机制建模分布值响应轨迹,结合基于稀疏度量学习的特征选择方法,提升预测准确性与可解释性,模拟及应用验证了其有效性。

中文摘要 AI 辅助

许多科学和工程应用产生的响应不是标量或向量,而是统计对象,其形式随时间、深度等有序索引演变。概率分布就是一个突出例子,它捕捉了低维统计无法概括的变异性和不确定性。当顺序观察此类响应时,动态分布轨迹给回归带来重大挑战。我们提出动态弗雷歇回归(DFR)框架,通过引入索引感知加权机制扩展全局弗雷歇回归。在每个索引处,预测定义为分布度量空间(如瓦瑟斯坦空间)中的加权弗雷歇均值。权重联合依赖于预测变量相似性和索引邻近性。为提高高维设置下的可解释性,DFR 结合基于稀疏度量学习的几何感知特征选择方法。模拟研究表明其预测准确性和特征恢复能力优于现有方法,增材制造数据应用展示了其产生可解释、特定索引分布预测的能力。

英文摘要

Many scientific and engineering applications generate responses that are not scalars or vectors, but statistical objects whose form evolves over an ordered index such as time, depth. Probability distributions are a prominent example, capturing variability and uncertainty that cannot be summarized by low-dimensional statistics. When such responses are observed sequentially, the resulting dynamic distributional trajectories pose significant challenges for regression, particularly in relating scalar predictors to both within-index variability and cross-index evolution. We propose Dynamic Fréchet Regression (DFR), a framework for modeling index-dependent trajectories of distribution-valued responses. DFR extends Global Fréchet Regression by introducing an index-aware weighting mechanism. At each index, predictions are defined as weighted Fréchet means in a metric space of distributions (e.g., Wasserstein space), preserving the intrinsic geometry of the response. The weights depend jointly on predictor similarity and index proximity, enabling index-specific prediction while borrowing strength across neighboring indices. To improve interpretability in high-dimensional settings, DFR incorporates a geometry-aware feature selection approach based on sparse metric learning, which identifies predictors driving distributional dynamics without relying on Euclidean coefficients. Simulation studies show improved predictive accuracy and feature recovery over existing methods. An application to additive manufacturing data demonstrates its ability to produce interpretable, index-specific distributional predictions.

↑