arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

将手工知识与学习表示融合何时有效?一种针对堆叠、替代和干扰的成本归一化基准

When does fusing hand-crafted knowledge with learned representations pay? A cost-normalized benchmark of stacking, substitution, and interference

Ahmad AlMughrabi, Albert Clop, Benjamin Busam, Ricardo Marques, Petia Radeva

arXiv 2608.21098首次发表:更新:

发表机构

Universitat de Barcelona; Technical University of Munich; Universitat Pompeu Fabra(巴塞罗那大学; 慕尼黑工业大学; 庞培法布拉大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究以成本归一化基准分析手工知识与学习表示融合的效果,发现三种结果并提出可预测端到端增益的分解式。

AI 中文摘要

在数据稀缺的场景中,将先验知识与数据驱动学习相融合具有吸引力,但目前尚无受控研究明确说明这种融合何时有益、冗余或有害。我们以一种固定的手工知识源(仅在训练期间注入的固定Gabor目标库,开销约为2%)为基准,对比数据驱动的替代方案(包括SimCLR、SimSiam、DINO、ImageNet迁移、数据增强、学习型教师),采用固定的实验方案和固定子集:13个数据集、9种骨干网络、图像规模150至128万张、分辨率32至224像素、参数规模250万至8600万,包含$\text{computeCells}$分类配置($\text{computeRuns}$次运行)及分割、检测移植任务。在我们测量的所有训练时间组合中,出现三种重复结果(决策级融合结果不同):不同“货币”的源可堆叠,该先验知识与DeiT增强在注意力骨干网络上结合,使ViT-B/16在224像素分辨率下性能提升26个点,预算翻倍时提升6.7个点;相同“货币”的源可替代,与有效的自监督预训练结合时,组合效果从未超过更优的单一源;将该知识完全融合到已初始化的模型中会产生干扰,干扰程度与知识携带量成正比:ImageNet迁移导致性能下降15至17个点,使用较弱的辅助权重可消除该影响。我们通过冻结特征诊断单独测量每个源,虽能回顾性区分这些结果,但无法预测:基于此构建的规则仅能在9个未见过的配对中正确预测1个。在从业者自身的标签预算下,冻结特征增益可在30个单元和7个数据集上将端到端增益预测误差控制在0.17个点以内,其底层分解式$\triangle = G + \text{readout}(\text{base})$在$\text{auditRate}\%$的可测试单元上符号成立,且可提前预测未见过的骨干网络族的特征增益。项目页面为该https URL。

英文摘要

Fusing prior knowledge with data-driven learning is attractive where data is scarce, yet no controlled account says when it helps, is redundant, or harms. We benchmark one fixed hand-crafted knowledge source, a pinned bank of Gabor targets injected only during training at $\sim$2\% overhead, against data-driven alternatives (SimCLR, SimSiam, DINO, ImageNet transfer, augmentation, learned teachers) under one frozen recipe with fixed subsets: 13 datasets, 9 backbones, 150 to 1.28M images, 32--224\,px, 2.5M--86M parameters ($\computeCells$ classification configurations over $\computeRuns$ runs, plus segmentation and detection transplants). Across the training-time combinations we measure, three outcomes recur (decision-level fusion differs). Different-\emph{currency} sources can stack: the prior composes with DeiT augmentation on attention backbones and is worth $+26$ points to ViT-B/16 at $224$\,px, $+6.7$ at twice that budget. Same-currency sources substitute: against effective self-supervised pretraining, the combination never usefully exceeds the better single source. Fusing at full strength into an already-informed initialization interferes in proportion to what it carries: ImageNet transfer, $-15$ to $-17$ points, removed by a weaker auxiliary weight. Frozen-feature diagnostics measured on each source alone separate these outcomes retrospectively but do not predict them: a rule built on them calls one of nine unseen pairs. At a practitioner's own label budget, the frozen-feature gain predicts the end-to-end gain to within $0.17$ points across 30 cells and seven datasets; the underlying decomposition, $Δ= G + \readout(\mathrm{base})$, holds in sign on $\auditRate\%$ of testable cells and is called an unseen backbone family's feature gain in advance. The project page is https://amughrabi.github.io/MomentAux.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑