哪一个印度在翻译?大型语言模型中印度口头传统的叙事同质化
Which India Survives Translation? Narrative Homogenisation Across Indian Oral Traditions in LLMs
浏览论文内容
中文总结 AI 辅助
本研究通过分析Claude Sonnet和Gemini两款LLMs对印度三种不同口头传统的文本生成,发现存在部分叙事同质化,且区域语言提示会降低传统文本保真度,为相关文化误解研究提供了轻量补充方法。
中文摘要 AI 辅助
大型语言模型(LLMs)主要基于过度代表某些文化叙事的英语互联网文本进行训练,引发了人们的担忧:模型会将非西方叙事传统的多样性压缩为单一同质化原型。我们开展了一项试点计算研究,在三种差异最大的印度区域口头与文学传统中考察这一问题:拉贾斯坦邦的帕布吉史诗、古典泰米尔桑盖姆诗歌以及孟加拉民间故事。我们为每种传统收集了真实参考语料库,分别包含11段、21段和10段文本,并针对每种传统的三种提示类型(通用型、文化特定型、区域语言型)向两款LLMs(Claude Sonnet和Gemini)发起54次生成请求。我们利用Sentence-BERT嵌入和余弦相似度,测量参考漂移(输出相对于另外两种传统,在多大程度上贴合自身传统的真实文本)与跨传统收敛性(输出在不同传统间的相似程度)。研究发现,尽管输出相对于其他传统更贴合自身传统的参考文本,但跨传统相似度较高(0.52至0.66),与这些传统真实差异所预示的水平相比,表明存在部分同质化现象。出乎意料的是,使用区域语言(印地语、泰米尔语或孟加拉语)提示时,与英语提示相比,对真实传统的保真度始终降低,其中拉贾斯坦邦和孟加拉传统的保真度降幅高达27个百分点。我们结合关于多语言提示的相互矛盾的先前结果对此展开讨论,认为这反映出在唤起一般文化多样性与模拟一种狭窄、文献记载较少的口头传统之间存在差异。我们将这项试点研究定位为近期关于LLM生成故事中印度文化误解的大规模人工标注研究的轻量、可扩展补充,是更广泛博士研究计划的一部分。
英文摘要
Large language models (LLMs) are trained predominantly on English-language internet text that over-represents certain cultural narratives, raising concerns that models flatten the diversity of non-Western storytelling traditions into a single homogenized archetype. We present a pilot computational study examining this across three maximally distinct Indian regional oral and literary traditions: the Rajasthani Pabuji epic, classical Tamil Sangam poetry, and Bengali folk tales. We collected authentic reference corpora for each tradition (11, 21, and 10 passages respectively) and prompted two LLMs (Claude Sonnet and Gemini) with 54 generation requests spanning three prompt types per tradition - generic, culturally specific, and regional-language. Using Sentence-BERT embeddings and cosine similarity, we measure reference drift (how closely outputs track their own tradition's authentic texts relative to the other two) and cross-tradition convergence (how similar outputs are across traditions). We find that while outputs remain closer to their own tradition's reference than to others, cross-tradition similarity is high (0.52-0.66) relative to what the traditions' genuine distance would predict, indicating partial homogenisation. Unexpectedly, prompting in the regional language (Hindi, Tamil, or Bengali) consistently reduced fidelity to the authentic tradition relative to English prompting, by as much as 27 percentage points for Rajasthani and Bengali traditions. We discuss this against conflicting prior results on multilingual prompting and argue it reflects a difference between eliciting general cultural diversity and simulating one narrow, lesser-documented oral tradition. We position this pilot as a lightweight, scalable complement to recent large-scale human-annotation studies of Indian cultural misrepresentation in LLM-generated stories, as part of a broader doctoral research program.