发表机构
Institute for People-Centered AI, University of Surrey; University of New South Wales(萨里大学以人为本人工智能研究所; 新南威尔士大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究英语方言适应中稳健性与生成的差距,通过DiaLLM持续预训练并结合多种策略对比不同英语变体。发现二者分离,特定变体适应产出受青睐但优化奖励方法未获评估者偏爱,缩小差距需丰富奖励设计和投入方言资源。
AI 中文摘要
大语言模型越来越能理解方言英语,但仍只能生成标准的、偏向美式的英语,方言生成这一难题基本未得到解决。我们引入了DiaLLM,它在国际英语语料库上对三个开放权重语言模型家族进行持续预训练,并应用隐式和显式的训练后范式,每种范式结合三种模型对齐策略,首次对澳大利亚、印度和英国北部英语的这些组件进行了对照比较。结果表明方言稳健性和生成是分离的:基准由持续预训练和监督微调塑造,而对齐以基准未捕捉的方式明显重塑生成。显式的针对特定变体的适应产生了被可靠识别为方言且比广泛对齐更受青睐的输出,但最积极优化方言奖励的方法并不受人类评估者青睐。独立的语言分析证实了这种奖励-质量差距。没有单一的对齐方法占主导,缩小差距需要更丰富的奖励设计和对方言资源的持续投入。我们发布了所有代码、检查点和偏好数据集。
英文摘要
Large language models increasingly understand dialectal English, yet still produce only standard, US-leaning English, leaving dialectal generation, the harder half of the problem, largely unaddressed. We introduce DiaLLM, which continually pretrains three open-weight language model families on the International Corpus of English and applies implicit and explicit post-training paradigms, each combined with three model alignment strategies, giving the first controlled comparison of these components across Australian, Indian, and Northern British English. Our results reveal a robustness-generation gap: benchmarks are shaped by continual pretraining and SFT, while alignment visibly reshapes generation in ways benchmarks do not capture. Explicit variety-targeted adaptation produces output reliably recognised as dialectal and judged more dialectal than broad alignment, yet where human judgement was directly assessed, the method that most aggressively optimises the dialectal reward is not the one judged most dialectal. Independent linguistic analysis corroborates this reward-quality gap, most clearly on two of the three families. No single alignment method dominates, and closing the gap will require richer reward designs and continued investment in dialectal resources. We release all code, checkpoints, and preference datasets.