发表机构
Blockchain Data Intelligence Lab; The Pennsylvania State University(区块链数据智能实验室; 宾夕法尼亚州立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过对比生成与选择任务,揭示LLM在文化特定亲属称谓上识别强但生成弱,提出需采用生成式评估基准以更全面衡量其能力。
AI 中文摘要
当前文献使用多项选择基准评估大型语言模型(LLM)在多语言亲属关系理解上的表现,将其视为一个识别问题。我们转而提示五个开放权重LLM在三种非西方语言(印地语、泰米尔语和韩语)中执行两项交际任务以生成亲属称谓,并将其与匹配的选项支持选择基线进行对比。在相同的亲属关系语言单元上,GPT OSS120B在75个有效单元中正确选择术语的比例为90.67%,但在相应尝试中产出可接受术语的比例仅为36.00%;Llama 3.370B表现出相同模式(77.92%对24.24%)。由于四选项条件展示了候选术语且不要求文字产出,这一差异被解释为评估格式差距,而非词汇知识完好的直接证据。在明确指定的L3提示下,准确率差异显著,从GLM-5.1的72.29%到Llama-3.370B的24.24%。父系谱系优势具有语言特异性;在印地语中显著,在韩语中较弱或反向,而泰米尔语共享术语对为测量变异提供了对照。这些结果表明,即使关系被明确陈述,文化特定亲属称谓的生成仍然困难,并促使在多项选择测试之外采用基于生成的评估。
英文摘要
Current literature evaluates large language models (LLMs) on multilingual kinship understanding using multiple choice benchmarks, treating it as a recognition problem. We instead prompt five open weight LLMs to generate kinship terms in three non Western languages (Hindi, Tamil, and Korean) across two communicative tasks and pair this with a matched option-supported selection baseline. On identical relation language cells, GPT OSS120B selects the correct term in 90.67% of 75 valid cells but produces an accepted term in 36.00% of the corresponding attempts; Llama 3.370B shows the same pattern (77.92% versus 24.24%). Since the four-option condition displays the candidate terms and does not require script production, the difference is interpreted as an evaluation format gap rather than direct proof that lexical knowledge is intact. On explicitly specified L3 prompts, accuracy varies sharply, from GLM-5.1 at 72.29% to Llama-3.370B at 24.24%. The paternal-lineage advantage is language specific; it is large in Hindi but weak or reversed in Korean, while Tamil shared-term pairs provide a control for measurement variation. These results show that culturally specific kinship generation remains difficult even when the relationship is explicitly stated and motivate generation-based evaluation alongside multiple-choice testing.
CommentsAccepted at (ORACLE Workshop), EMNLP 2026