发表机构
Faculdade de Ciências e Tecnologia da Universidade do Algarve(阿尔加维大学科学与技术学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对近红外化学计量学中CNN架构设计问题,通过计算光谱数据集属性描述符,利用贝叶斯超参数优化优化两个一维CNN支架,得出光谱描述符可提供设计先验,引导浅层模型至合理超参区域,且相关启发式方法与HPO有竞争力。
AI 中文摘要
用于近红外(NIR)化学计量学的卷积神经网络(CNN)通常采用通用架构规则设计,尽管光谱数据集在采样、平滑度、冗余度和样本大小等方面存在差异。我们测试了这些属性是否能为CNN设计提供经验先验。在25个近红外回归任务中,计算了数据集大小、光谱长度和间距、熵、内在秩、自相关和小波尺度结构的描述符。使用五折交叉验证的贝叶斯超参数优化(HPO)对两个可解释的一维CNN支架(一个最小单卷积模型和一个带有可选分支、扩张等的扩展浅模型)进行了优化。从接近最优试验中提取的关系被转换为热启动启发式方法,并通过留一数据集法(LODO)验证直接评估。最明显的关系涉及卷积感受野。在最小CNN中,首选内核分数随光谱熵和内在秩的增加而减小,随小波能量支持分数的增加而增加,学习率倾向于随训练集大小的增加而减小。直接和LODO启发式方法与HPO具有竞争力,中位数测试RMSE比率分别为0.953和1.017。扩展CNN在分支使用、扩张、随机失活、滤波器数量和感受野选择方面显示出相似但不太可转移的结构。十次随机重新拟合显示出与HPO选择的配置相当的种子敏感性。在另一个实验中,联合预处理和CNN HPO在25个任务中的19个任务中优于标准化光谱HPO,尽管增益取决于数据集。这些结果表明,光谱描述符可以提供实用的CNN设计先验,在进行特定目标调整之前,将浅层近红外模型引导到合理的超参数区域。
英文摘要
Convolutional neural networks (CNN) for near-infrared (NIR) chemometrics are often designed using generic architectural rules, although spectral datasets differ in sampling, smoothness, redundancy, and sample size. We tested whether these properties can provide empirical priors for CNN design. Across 25 NIR regression tasks, we computed descriptors of dataset size, spectral length and spacing, entropy, intrinsic rank, autocorrelation, and wavelet-scale structure. Two interpretable 1D-CNN scaffolds (a minimal single-convolution model and an extended shallow model with optional branching, dilation, etc) were optimized using five-fold cross-validated Bayesian hyperparameter optimization (HPO). Relationships extracted from near-optimal trials were converted into warm-start heuristics and evaluated directly and through leave-one-dataset-out (LODO) validation. The clearest relationships involved convolutional receptive fields. In the minimal CNN, the preferred kernel fraction decreased with spectral entropy and intrinsic rank, increased with the wavelet energy-support fraction, and the learning rate tended to decrease with training-set size. Direct and LODO heuristics were competitive with HPO, with median test-RMSE ratios of 0.953 and 1.017, respectively. The extended CNN showed similar but less transferable structure across branch usage, dilation, dropout, filter counts, and receptive-field choices. Ten stochastic refits showed seed sensitivity comparable to that of HPO-selected configurations. In a separate experiment, joint preprocessing and CNN HPO outperformed standardized-spectra HPO in 19 of 25 tasks, although gains were dataset-dependent. These results show that spectral descriptors can provide practical CNN design priors, guiding shallow NIR models toward plausible hyperparameter regions before target-specific tuning
Comments25 pages, 4 figures, submitted to peer-review journal