arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.18394cs.AI

情境中的文化能力:大型语言模型在芬兰通过图灵测试

Cultural Competence in Context: A Large Language Model Passes the Turing Test in Finland

Otto Segersven, Pentti Henttonen

首次发表
浏览论文内容

中文总结 AI 辅助

本研究在芬兰用芬兰语进行图灵测试,发现大型语言模型ChatGPT 5.2意外通过,并提出模型生成角色提示作为可复制的比较测试方法,将图灵测试重新定义为检验AI在特定社会世界中成员身份可信度的工具。

中文摘要 AI 辅助

我们报告了在芬兰以芬兰语进行的一项图灵测试的结果。由于语言和文化背景在LLM训练数据中的代表性不均,我们预期该模型(ChatGPT 5.2)在芬兰语图灵测试中的表现会不如先前研究的英语美国情境。我们还提出了模型生成的角色提示,作为一种可复制的技术,用于进行基于LLM的比较性图灵测试,旨在提高构念效度。与我们的预期相反,该LLM通过了芬兰语图灵测试。一个显著的错误来源是参与者依赖语言线索,尤其是口语化的芬兰语,作为人类作者身份的标记。我们将图灵测试从智力测试重新定义为一种比较方法,用于检验AI系统是否能在特定社会世界中展示可信的成员身份。由于其结果反映了模型能力、提示身份、人类参与者的内部能力及其AI素养,该方法为跨领域探索人机边界提供了有用的探针。

英文摘要

We report the results of a Turing Test conducted in Finland in the Finnish language. Because languages and cultural contexts are unevenly represented in LLM training data, we expected the model (ChatGPT 5.2) to perform worse in a Finnish-language Turing Test than in previously studied English-language US contexts. We also present model-generated role prompting as a replicable technique for conducting comparative LLM-based Turing Tests designed to improve construct validity. Contrary to our expectations, the LLM passed the Finnish Turing Test. A prominent source of error was participants' reliance on linguistic cues, particularly colloquial Finnish, as markers of human authorship. We reframe the Turing Test from a test of intelligence to a comparative method for examining whether an AI system can display credible membership in a particular social world. Because its outcome reflects model capabilities, prompted identity, insider competence among human participants, and their AI literacy, the method provides a useful probe of the human-machine boundary across domains.

发表机构

  • University of Helsinki(赫尔辛基大学)
  • University of Tampere(坦佩雷大学)

机构由 AI 辅助整理,请以论文原文为准。

↑