发表机构
Preply(普雷普利)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对大规模语法标注成本高的问题,微调 Qwen3.5 小型语言模型并部署 0.8B 模型,其性能优于提示式 GPT-5.4、GPT-5.6 Sol,还降低了服务成本并提升了学习者参与度与业务指标。
AI 中文摘要
纠正性反馈是二语习得最有证据支持的驱动因素之一,但课程中提供的纠正很少能积累成可操作的语法掌握视图。提示式前沿模型可从学习者-教师课程 transcript 中提供此类视图,但大规模应用成本高昂。我们通过在过滤和重新平衡的教师生成监督数据上微调 Qwen3.5 小型语言模型(SLMs),然后在我们平台上为所有英语学习者部署高效的 0.8B 模型,构建端到端语法掌握跟踪器,从而缩小这一差距。将标注契约内化到适配器权重中,使 0.8B 模型能与紧凑匹配的提示配对,而非冗长指令。在两个人工策划的基准测试中,部署的 0.8B 模型和 4B 参考比较器在嵌套匹配标准(严格性递增:概念、证据跨度、正确性)下的精确率和召回率均优于提示式 GPT-5.4 和 GPT-5.6 Sol。部署的 0.8B SLM 将服务成本降低约 16×。特征级在线实验显示学习者参与度显著提升(+15.8%),关键业务指标也有所改善,包括计划时长(+2.1%)和新课程 GMV(+13.2%)。
英文摘要
Corrective feedback is among the best-evidenced drivers of second-language acquisition, yet corrections delivered during lessons rarely accumulate into an actionable view of grammar mastery. Prompted frontier models can provide such a view from learner--tutor lesson transcripts, but they are costly at scale. We close this gap by fine-tuning Qwen3.5 small language models (SLMs) on filtered and rebalanced teacher-generated supervision, then deploying an efficient 0.8B model in an end-to-end grammar mastery tracker for all English learners on our platform. Internalizing the annotation contract into adapter weights enables pairing the 0.8B model with a compact matched prompt rather than verbose instructions. On two human-curated benchmarks, both the deployed 0.8B model and a 4B reference comparator outperform prompted GPT-5.4 and GPT-5.6 Sol in precision and recall under nested matching criteria of increasing strictness: concept, evidence span, and correctness. The deployed 0.8B SLM reduces serving cost by approximately 16$\times$. A feature-level online experiment shows significant gains in learner engagement ($+15.8\%$) and key business metrics, including scheduled hours ($+2.1\%$) and GMV from new lessons ($+13.2\%$).