arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基础模型在婴儿信号上的迁移能力如何?基于统一需求本体的跨数据集迁移审计

How Far Do Foundation Models Transfer to Infant Signals? A Cross-Dataset Transfer Audit with a Unified Need Ontology

Wu Hangyu

arXiv 2608.08989首次发表:更新:

发表机构

Shenzhen Coddie Technology Co., Ltd.(深圳科迪科技有限公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对婴儿哭声语料库的标签不兼容等问题,通过多级审计探究基础模型的跨数据集迁移能力,提出实用方案并发布工具将不兼容语料库转为联合训练资源。

AI 中文摘要

公开的婴儿哭声语料库规模小、标签不兼容,且几乎总是仅针对单个语料库进行评估。我们探究这种做法隐藏了什么问题以及有什么解决方案。我们通过多级泄露审计(字节级和嵌入级去重,加上语料库内训练-测试近重复审计)筛选了四个哭声语料库,在统一的五类需求本体和共享任务设定下,对四个冻结编码器和一个手工基线进行了探究。该审计揭示了单语料库评估所掩盖的问题:同一编码器的域内宏F1分数波动达0.57-0.80;跨语料库迁移平均为负(负迁移率为0.19-0.35,在30个定向单元中有18个显著,经BH-FDR校正);349个内容相同的片段组在不同语料库分布中存在矛盾的元数据标签。然而,该审计也揭示了可行的改进方向:在匹配训练规模并移除近重复后,向最嘈杂语料库的迁移效果始终为正,为小型嘈杂语料库提供了实用方案;冻结探针在适度标签预算下达到饱和,而稳定微调在全标签设置下表现更优;域自适应预训练在5-10次少样本设置下显著优于稳定微调(1次少样本优势对优化种子方差不稳健),但在50次少样本及以上设置下无显著优势;在测试的二元共享标签设定中,本体映射联合训练在所有四个编码器-目标组合中均胜出,而直接合并未映射标签会导致F1分数损失高达37分。我们发布了该本体、映射代码和审计流程,将不兼容的哭声语料库转化为可用的联合训练资源。

英文摘要

Public infant cry corpora are small, label-incompatible, and almost always evaluated one corpus at a time. We ask what this practice hides and what fixes it. Across four cry corpora screened by a multi-level leakage audit (byte-level and embedding-level deduplication plus a within-corpus train-test near-duplicate audit), we probe four frozen encoders and a handcrafted baseline under a unified five-class need ontology and shared task formulations. The audit exposes what single-corpus evaluation conceals: within-domain macro-F1 swings by 0.57-0.80 for the same encoder, cross-corpus transfer is negative on average (negative-transfer ratio 0.19-0.35, significant in 18 of 30 directed cells, BH-FDR), and 349 content-identical clip groups carry conflicting metadata labels across corpus distributions. The same audit, however, reveals a consistent way forward. Transfer into the noisiest corpus is consistently positive in effect size at matched training size and after near-duplicate removal, offering a practical recipe for small, noisy corpora. Frozen probes saturate at modest label budgets, while stabilized fine-tuning wins with full labels; domain-adaptive pretraining significantly beats stabilized fine-tuning at 5-10-shot (the 1-shot advantage is not robust to optimization-seed variance) but shows no significant advantage at 50-shot or beyond. In the tested binary, shared-label settings, ontology-mapped joint training wins in all four encoder-by-target combinations, whereas naively merging unmapped labels costs up to 37 F1 points. We release the ontology, mapping code, and audit pipeline, turning incompatible cry corpora into a usable joint-training resource.

Comments18 pages, 7 figures. Under review at AAAI 2027

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑