发表机构
National Yang Ming Chiao Tung University(国立阳明交通大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究设备端英到繁体中字幕翻译,针对短输入等条件限制,通过替换词汇表、迁移嵌入空间及校准微调模型,在OpenSubtitles2024测试集上有较好胜率,苹果M2 Metal测量有加速,为实时字幕翻译提供优化方法。
AI 中文摘要
本报告研究在短输入、短输出、单批次推理、低延迟和隐私约束下,针对台湾地区的设备端英文到繁体中文的字幕翻译。这些条件限制了为长上下文或高吞吐量语言模型服务设计的优化价值。从LMT - 60 - 0.6B开始,初步分析表明在GGUF量化降低了Transformer块的相对成本后,词汇投影成为解码时更重要的成本。我们用64k令牌的字幕域分词器替换了原来的151k令牌词汇表,迁移了嵌入空间,并通过嵌入校准和全监督微调来调整模型。在OpenSubtitles2024测试集的固定500个示例子集上,LocalSubs在GPT - 4o成对判断下相对于谷歌翻译实现了59.2%的排除平局胜率。在短提示上性能最强,随着提示长度增加而下降。在64k词汇模型上的初步苹果M2 Metal测量显示,相对于151k词汇分析基线有1.63倍的加速。原始基准配置不完整,因此延迟结果视为初步结果。
英文摘要
This report studies on-device English-to-Traditional-Chinese subtitle translation for Taiwan under short inputs, short outputs, batch-size-one inference, low latency, and privacy constraints. These conditions limit the value of optimizations designed for long-context or high-throughput language-model serving. Starting from LMT-60-0.6B, preliminary profiling suggests that vocabulary projection becomes a more important decode-time cost after GGUF quantization reduces the relative cost of Transformer blocks. We replace the original 151k-token vocabulary with a 64k-token subtitle-domain tokenizer, migrate the embedding space, and adapt the model through embedding calibration followed by full supervised fine-tuning. On an OpenSubtitles2024 test set, LocalSubs achieves a 59.2% tie-excluded win rate against Google Translate under GPT-4o pairwise judging. Performance is strongest on short cues and declines as cue length increases. In a separate preliminary Apple M2 Metal profiling run, LocalSubs shows a 1.63x speedup over a 151k-vocabulary baseline. The code is available on https://github.com/aiden1020/localsubs .