迷失在翻译中:衡量非母语英语对大型语言模型终端用户表现的影响
Lost in Translation: Measuring the Effect of Non-Native English on End User Performance of Large Language Models
查看机构详情
- University of Michigan(密歇根大学)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
本研究通过构建FABLE数据集,评估34个开源大语言模型对非母语英语提示的响应,发现模型虽不传播拼写错误,但会镜像用户的修辞和词汇质量,导致非母语用户获得质量更低、流利度更差的答案。
中文摘要 AI 辅助
大型语言模型(LLMs)越来越多地被母语非英语的人群使用,然而研究表明,这些用户收到的回复质量系统性地低于流利英语使用者。究竟是哪些非母语英语的具体特征导致了这一差距仍不清楚,因为流利度本身是由机械准确性、词汇使用、组织结构和语篇连贯性共同构成的复合体。在此,我们引入了FABLE,一个受控数据集,包含190,911个英语提示变体,这些变体源自174K个真实用户针对写作相关任务的提示。通过评估34个开放权重模型(open-weight LLMs)的回复,我们发现了一个明显的非对称性:虽然模型不会将拼写错误等表面错误传播到其输出中,但模型确实会镜像用户提示中存在的更高层次的修辞和词汇质量。此外,在最不流利和最流利的提示之间,回复的整体质量存在显著差异。这些结果凸显了非母语英语LLM用户所面临的关键性能差距,导致他们获得质量更低且流利度更差的答案。
英文摘要
Large language models (LLMs) are increasingly used by people whose first language is not English, yet these users have been shown to receive systematically lower-quality responses than fluent speakers. Which specific features of non-native English drive this gap remains unclear, because fluency is itself a composite of mechanical accuracy, vocabulary use, organization, and discourse coherence. Here, we introduce FABLE, a controlled dataset of 190,911 English prompt variants derived from 174K real user prompts for writing-related tasks. Evaluating responses from 34 open-weight LLMs, we find a clear asymmetry; while models do not propagate surface errors such as misspellings into their outputs, models do mirror higher-level rhetorical and lexical qualities present in the user's prompt. Further, the overall quality of responses differs substantially between the least- and most-fluent prompts. These results highlight a key LLM performance disparity for non-native English LLM users, resulting in both lower-quality and less-fluent answers.