arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

语音到SFT流水线的因子消融:对数据质量与下游迁移的差异化影响

A Factorial Ablation of a Speech-to-SFT Pipeline: Differential Effects on Data Quality and Downstream Transfer

Wonsup Shin, Jingu Kim

arXiv 2608.20394首次发表:更新:

发表机构

Flitto(飞力拓)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究对语音转SFT流水线进行因子消融,发现QA数据质量提升未均匀转化为下游MCQA增益,正向迁移集中于家族领域匹配配对,验证了流水线稳健性并发布相关资源。

AI 中文摘要

将语音经多阶段优化转化为监督微调(SFT)数据的工业流水线正被越来越多地采用,但据我们所知,这些流水线尚未公开进行逐阶段消融分析,导致各阶段的边际价值尚不明确。我们设计了一款可投入实际使用的语音到SFT流水线,其中转录优化(第0阶段)与SFT数据质量优化(第2阶段)可独立切换,形成2×2因子设计。针对每种情况,我们从韩语医疗与金融会议录音中生成问答(QA)格式的SFT数据,并微调9个模型(5个大语言模型(LLM)家族,参数规模24亿至700亿);我们采用4个跨提供商LLM评判器、盲法6名专家人工评估及3个下游多项选择问答(MCQA)基准进行评估。核心发现:在固定且标准的SFT配方下,QA数据质量的提升并不会均匀转化为下游MCQA收益;4个评判器的质量评分持续上升,但跨模型的平均MCQA增益并不显著;正向迁移集中在家族与领域匹配的配对中。这种差异化模式与格式不匹配一致:第2阶段将SFT数据构成转向解释性条目,而MCQA主要探查事实性回忆。所有6名人形评判者均报告全流水线质量更高,证实了LLM评判器的方向。将语音识别(STT)引擎替换为Whisper-medium后,验证了流水线的稳健性。非幻觉审计显示,两个前沿LLM平均在约8%的QA问题上承认未知;我们发布了样本、提示、代码及所有SFT检查点。

英文摘要

Industry pipelines that turn speech into supervised fine-tuning (SFT) data via multi-stage refinement are increasingly adopted but, to our knowledge, have not been publicly ablated stage-by-stage, leaving each stage's marginal value unknown. We design a production-ready speech-to-SFT pipeline in which transcript refinement (Phase 0) and SFT data quality refinement (Phase 2) are independently toggleable, yielding a 2x2 factorial design. For each condition, we generate QA-form SFT data from Korean medical and finance conference recordings and fine-tune 9 models (5 LLM families, 2.4B-70B); we evaluate with four cross-provider LLM judges, a blind six-expert human evaluation, and 3 downstream MCQA benchmarks. Our central finding: under a fixed, standard SFT recipe, improvements in QA data quality do not transfer uniformly into downstream MCQA gains. 4-judge quality rises consistently, yet the cross-model mean MCQA gain is not significant; positive transfer concentrates on family-domain aligned pairs. This differential pattern is consistent with a format mismatch: Phase 2 shifts SFT-data composition toward explanatory items, while MCQA primarily probes factoid recall. All six human raters report higher full-pipeline quality, confirming the LLM-judge direction. An STT-engine swap to Whisper-medium confirms pipeline robustness. A non-hallucination audit shows the two frontier LLMs admit unknown on approximately 8% of QA on average; we release samples, prompts, code, and all SFT checkpoints.

Comments20 pages, 2 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑