arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CHARM:面向跨文化角色扮演基准的角色幻觉研究

CHARM: Character Hallucination for Multicultural Role Play Benchmark

Sunkyung Han, Nahyeon Park, Gaeun Seo, Seunghyun Yoon, JinYeong Bak

arXiv 2609.01352首次发表:更新:

发表机构

Sungkyunkwan University; Adobe Research(成均馆大学; 奥多比研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出跨文化角色扮演基准CHARM,通过两阶段评估区分大语言模型的边界意识与遵守能力,发现角色幻觉主要由遵守失败驱动,且存在系统性文化差异。

AI 中文摘要

角色扮演类大语言模型(LLM)需同时贴合角色风格并尊重角色的知识边界。现有评估虽能检测角色幻觉,但极少区分错误源于无法识别边界,还是识别后仍不遵守。本文提出CHARM,这是一个涵盖5个文化语言区域的40个真实与虚构角色的跨文化基准,且经母语评审验证。它采用支持弃权(不执行)的多项选择题,探测时间(历史与现代)和跨宇宙(角色叙事或历史宇宙之外的实体)两类边界。本文提出两阶段评估,将边界意识(明确识别查询超出范围)与边界遵守(回答具体问题时弃权)分开。对6个LLM的评估显示,幻觉主要由遵守失败驱动:模型常承认查询超出角色知识,却仍提供角色外的事实性答案。通过向目标角色重提相同问题,本文确认大部分此类情况是经参数覆盖的;模型存储了相关事实,但无法抑制它。本文还观察到这些失败存在系统性文化差异,与模型知识中不同区域角色的表征失衡一致。

英文摘要

Role-playing large language models (LLMs) are expected to adopt a character's style while also respecting that character's knowledge boundaries. Prior evaluations detect character hallucination but rarely distinguish whether errors arise from failure to recognize a boundary or from failure to comply despite recognition. We introduce CHARM, a multicultural benchmark of 40 real and fictional characters drawn from five cultural-linguistic regions, and validated by native reviewers. It probes two boundary types, Temporal (historical vs. modern) and Cross-Universe (entities outside a character's narrative or historical universe), using abstention-enabled multiple-choice questions. We propose a two-stage evaluation that separates Boundary-Awareness (explicit recognition that a query is out of scope) from Boundary-Compliance (abstention when answering concrete questions). Evaluations across six LLMs show that hallucination is driven predominantly by compliance failures. Models frequently acknowledge that a query lies outside the character's knowledge yet still provide factual, out-of-character answers. By re-posing the same questions to the target character, we confirm that a large fraction of these cases are verified parametric overrides; the model stores the relevant fact but fails to suppress it. We also observe systematic cultural variation in these failures, consistent with imbalances in how characters from different regions are represented in model knowledge.

Comments16 pages, 1 figure. Accepted to Findings of EMNLP 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑