归因于光谱基础模型中的预处理不变性
Attributing Preprocessing Invariance in Spectral Foundation Models
浏览论文内容
中文总结 AI 辅助
该研究以拉曼光谱基础模型为例,指出其声称的预处理不变性多源于归一化而非学习,多数系统的归一化已消除相关变换,训练编码器未带来显著增益。
中文摘要 AI 辅助
预处理不变性是光谱基础模型的一个理想目标:当实验室对光谱进行不同的预处理时,冻结的模型应仍然可用。通常通过在一种预处理流程下训练分类器,并在另一种预处理流程下对其进行测试来衡量这种不变性,而保留的准确率被视为学习有效的证据。我们以拉曼光谱基础模型为例重新审视这一解读。这类模型在应用任何学习到的参数之前会对输入进行归一化处理。如果该归一化将两种不同预处理的光谱映射到同一向量,编码器将接收到相同的输入,因此这种不变性不能归因于学习。对于使用每个光谱自身统计量的归一化,当一个光谱是另一个光谱的正倍数加上一个常数时,就会发生这种情况。几种标准的预处理操作都呈现这种形式。因此,编码器应仅针对归一化进行评估,而归一化没有学习参数。在六个拉曼评估数据集上,该模型的表现并未显著优于其自身的归一化。它比原始光谱有所改进,但仅归一化也能做到这一点。训练确实使编码器在随机初始化的基础上有所提升,且一项受控实验表明,只有当某个变换作用于编码器时,它才会学习忽略该变换。一项数值测试确定了给定归一化会消除哪些变换。在五种模态的已发布系统中,大多数归一化已消除了该形式的变换,且其中几个系统声称将这种不变性作为学习成果。对其中两个系统重复该比较也显示没有增益。
英文摘要
A spectral foundation model should remain useful when laboratories preprocess spectra differently. The standard test trains a classifier under one pipeline and evaluates under another, taking preserved accuracy as evidence of learned invariance. However, these models normalize each input before any learned parameter is applied. When normalization maps differently preprocessed spectra to the same vector, the encoder receives identical inputs and the measured invariance cannot be attributed to learning. We propose a normalization-only attribution control: compare the encoder against its normalization before interpreting transfer as learned invariance. On six Raman datasets the encoder does not measurably improve transfer over its normalization. A controlled experiment confirms invariance develops only when variation reaches the encoder past normalization. Across three systems, no encoder improves relative retention over its normalization. An audit of eighteen configurations across five modalities confirms the issue is widespread: the normalization-only control should be reported before crediting transfer to the encoder.
发表机构
- ESCP Business School(ESCP商学院)
- OpluxCare Co., Ltd.(欧普乐康护理有限公司)
- The University of Hong Kong(香港大学)
- Nanjing University(南京大学)
机构由 AI 辅助整理,请以论文原文为准。