arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

冻结嵌入生物声学分类中域对齐的强度单调定律

A Strength-Monotonic Law for Domain Alignment in Frozen-Embedding Bioacoustic Classification

Yucheng Gong, Rui Zhou, Binbin Zeng, Qiang Ren, Hongjin Hui

arXiv 2610.09737首次发表:更新:

发表机构

Chongqing Finance and Economics College(重庆财经学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究揭示编码器越强越依赖MMD对齐且越受重平衡损害,据此简化配方为冻结嵌入加单MMD,保持性能并提升未见域泛化。

AI 中文摘要

分布对齐何时能帮助冻结的基础模型嵌入在声学域间泛化?针对跨域蚊子物种分类,我们报告了一个强度单调定律:编码器在目标任务上越强,其未见域泛化越依赖于分布对齐(MMD)项,且越受域重平衡采样的损害。在四个编码器家族和编码器内HuBERT层扫描(n=8)中,重平衡分支与编码器强度完全有序(Spearman -1.000),而MMD收益分支在每个流内单调,合并后为-0.857;固定架构仅改变表示强度会将重平衡效应从收益翻转为崩溃。该定律具有可操作性:单个MMD项是强编码器上的唯一杠杆,因此我们将该领域的默认配方简化为冻结的Perch 2.0嵌入、轻量探针、交叉熵、一个MMD和输入增强。简化配方保持在完整组合的种子噪声范围内(BA_unseen 0.299±0.006 vs. 0.307±0.014)。作为同一定律的边界条件,三个社区默认设置(骨干微调、多模态融合和域重平衡)在留域协议下均损害未见域准确率,通过单变量、多种子证据展示。我们提出一种机制及其解释的配方,而非排行榜条目。

英文摘要

When does distribution alignment help a frozen foundation-model embedding generalize across acoustic domains? For cross-domain mosquito-species classification we report a strength-monotonic law: the stronger an encoder is on the target task, the more its unseen-domain generalization relies on a distribution-alignment (MMD) term, and the more it is harmed by domain-rebalanced sampling. Across four encoder families and a within-encoder HuBERT layer sweep (n=8), the rebalancing leg orders exactly with encoder strength (Spearman -1.000), while the MMD-benefit leg is monotonic within each stream and -0.857 pooled; fixing architecture and varying only representation strength flips the rebalancing effect from benefit to collapse. The law is actionable: a single MMD term is the sole lever on a strong encoder, so we reduce the field's default recipe to a frozen Perch 2.0 embedding, a lightweight probe, cross-entropy, one MMD, and input augmentation. The reduced recipe stays within seed noise of the full composite (BA_unseen 0.299+/-0.006 vs. 0.307+/-0.014). As boundary conditions of the same law, three community defaults (backbone fine-tuning, multi-modal fusion, and domain rebalancing) each hurt unseen-domain accuracy under a leave-domain protocol, shown with single-variable, multi-seed evidence. We present a mechanism and the recipe it explains, not a leaderboard entry.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑