发表机构
Blossom AI Labs(Blossom AI 实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过合成数据流程和LoRA参数高效适配,使紧凑语音模型在日语护理交接任务上获得接近全量微调的性能,展示了从少量合成示例获取窄域音频到结构转换能力的可行性。
AI 中文摘要
私有领域语音难以收集和再分发,而紧凑模型需要任务特定的监督。我们研究了一种可审计的合成流程,将日语护理交接直接映射为六字段结构化笔记。使用182个合成训练和开发片段,我们通过全量微调和秩16的LoRA适配了一个1.47B的音频模型。在39个片段、场景和种子不重叠的合成测试集上,未适配模型获得模型评判的事实性召回得分为0.0500,全量微调为0.8664,LoRA为0.8461。LoRA达到全量微调总分的97.7%(作为描述性比率),而项目遥测报告可训练参数为1240万,约占骨干网络的0.85%。两种适配相对于同一基础模型均显示出较大的成对增益;全量与LoRA的区间跨越零,且不同的优化设置排除了等价性声明。这是一项参数高效的能力获取结果,而非设备性能结果:延迟、内存、能耗和实时因子均未测量。所有评估语音和目标均为合成数据,参考由模型提出,评判器未校准。证据表明,紧凑模型可以从数百个来源关联的合成示例中获取窄范围的音频到结构转换能力;但这并不确立临床有效性、真实语音迁移或相对于干净云系统的优越性。
英文摘要
Private domain speech is difficult to collect and redistribute, while compact models need task-specific supervision. We study an auditable synthetic pipeline that maps Japanese care handoffs directly to six-field structured notes. Using 182 synthetic training and development clips, we adapt a 1.47B audio model by full fine-tuning and rank-16 LoRA. On a 39-clip scenario-seed-disjoint synthetic test, an unadapted model obtains a model-judged factuality-recall score of 0.0500, full tuning 0.8664, and LoRA 0.8461. LoRA reaches 97.7% of the full-tuning aggregate as a descriptive ratio while project telemetry reports 12.4M trainable parameters, about 0.85% of the backbone. Both adaptations show large paired gains over the same base; the full-versus-LoRA interval crosses zero, and differing optimization settings preclude an equivalence claim. This is a parameter-efficient capability-acquisition result, not a device-performance result: latency, memory, energy, and real-time factor were not measured. All evaluation speech and targets are synthetic, references are model-proposed, and the judge is uncalibrated. The evidence shows that a compact model can acquire a narrow audio-to-structure transformation from a few hundred provenance-linked synthetic examples; it does not establish clinical validity, real-speech transfer, or superiority to a clean cloud system.
CommentsSubmitted to On-Device Intelligence: Foundation Models under Real-World Constraints (NeurIPS 2026 workshop). 4 pages, 0 figures, 1 table