arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.04393cs.LG

当代理预测变为方程重构:因子衍生代理监督的诊断与残差学习

When Proxy Prediction Becomes Equation Reconstruction: Diagnostics and Residual Learning for Factor-Derived Proxy Supervision

Chayan Lahiri, Ahmed Shafee, Cody Fehringer

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对科学机器学习中因子衍生代理预测的方程重构问题,提出RASPL公式保留残差框架,在RUSLE土壤流失预测任务中显著提升了退化与尾部鲁棒性。

中文摘要 AI 辅助

当直接观测有限时,科学机器学习常依赖从已知领域因子计算得到的代理目标;然而,当这些相同因子被用作模型输入时,高预测精度可能反映的是代理生成方程的重构,而非对退化因子信息的鲁棒性。我们在基于RUSLE(通用土壤流失方程)的土壤流失代理预测中,对土壤可蚀性因子K进行受控退化,以研究该问题。我们引入了一套诊断框架,其结合了退化公式参考、经典基于树的基线、匹配的直接与公式特征预测器、上下文消融、尾部误差分析以及退化鲁棒性评分。随后,我们提出了RASPL(公式保留残差框架),该框架将退化公式估计值保留为预测锚点,并学习自适应门控的上下文修正。RASPL的性能显著优于匹配的直接预测,且相较于将公式估计值视为普通输入特征的方法,其具备更强的退化鲁棒性和尾部鲁棒性。在RASPL内部,紧凑统计编码器实现了最高的宏平均R²和最低的计算成本,而卷积编码器则实现了最强的退化鲁棒性和最低的Tail95平均绝对误差(MAE)。这些结果确立了公式保留作为从因子衍生代理目标进行鲁棒学习的核心设计原则。

英文摘要

Scientific machine learning often relies on proxy targets computed from known domain factors when direct observations are limited. When those same factors are used as model inputs, however, high predictive accuracy may reflect reconstruction of the proxy-generating equation rather than robustness to degraded factor information. We study this problem in RUSLE-derived soil-loss proxy prediction under controlled degradation of the soil-erodibility factor $K$. We introduce a diagnostic framework that combines degraded-formula references, classical tree-based baselines, matched direct and formula-feature predictors, contextual ablations, tail-error analysis, and degradation robustness scoring. We then propose RASPL, a formula-preserving residual framework that retains the degraded formula estimate as the prediction anchor and learns an adaptively gated contextual correction. RASPL substantially outperforms matched direct prediction and provides stronger degradation and tail robustness than treating the formula estimate as an ordinary input feature. Within RASPL, a compact statistical encoder achieves the highest macro-averaged $R^2$ and lowest computational cost, whereas a convolutional encoder achieves the strongest degradation robustness and lowest Tail95 mean absolute error (MAE). These results establish formula preservation as the central design principle for robust learning from factor-derived proxy targets.

补充信息

↑