arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.01629cs.CL

在非母语日语的语言态度上实现人类与大语言模型(LLM)的对齐

Human-LLM Alignment in Language Attitudes Toward Non-Native Japanese

Naho Orita, Hayato Ogawa, Daisuke Kawahara

AI总结:

该研究对比人类与LLM对母语及非母语日语邮件的评价,发现LLM以结构化弱化形式重现母语者的语言态度,为审计LLM提供了可推广的语言态度框架。

AI中文摘要:

大语言模型(LLM)正越来越多地在招聘、学术评估等高风险领域对人类写作进行评价,这使得非母语使用者面临特殊风险。本研究采用语言态度框架,将人类与LLM对平行的母语(L1)及第二语言(L2)日语邮件的评价,在流畅度、地位、团结度三个维度进行对比。日语母语评分者在所有三个维度上对第二语言文本的评分显著更低,其中流畅度差距约为地位与团结度差距的两倍。六个LLM评判者重现了该偏差的方向,五个重现了维度间的排序。模型与人类的分歧体现在两方面:所有模型均低估了最具社会属性的团结度差距,且所有模型都区分了学习者的母语背景,而人类未做此区分。因此,LLM评判者以结构化但弱化的形式重现了母语使用者的语言态度,语言态度框架为审计LLM(不限于英语领域)提供了现成的标准。

英文摘要:

Large language models (LLMs) increasingly evaluate human writing in high-stakes domains such as hiring and academic assessment, putting non-native speakers at particular risk. Drawing on the language attitudes framework, we compared human and LLM evaluations of parallel L1- and L2-written Japanese emails on three dimensions: fluency, status, and solidarity. Japanese raters rated L2 texts significantly lower on all three dimensions, with a fluency gap roughly twice the size of the status and solidarity gaps. Six LLM judges reproduced the direction of this bias, and five reproduced its ordering across dimensions. The models diverged from humans in two ways: all understated the solidarity gap, the most socially grounded dimension, and all differentiated among learner L1 backgrounds where humans did not. LLM judges thus reproduce native speakers' language attitudes in a structured yet attenuated form, and the language attitudes framework offers a ready-made yardstick for auditing them beyond English.

↑