arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.05576cs.CLcs.CY

模型收敛与人类分歧:开放式生成中分布多元性的覆盖框架

Where Models Converge and Humans Diverge: A Coverage Framework for Distributional Pluralism in Open-Ended Generation

Zini Yang, Emily Wenger, Richard So

首次发表
浏览论文内容

中文总结 AI 辅助

本文针对LLM与人类写作间的分布差距问题,提出以人类写作经验分布为基准的框架,通过LLM-Cov和IBR指标测量LLM生成内容的分布广度,发现当前LLM内容合理但狭窄,可用于评估其“文化覆盖范围”。

中文摘要 AI 辅助

当大型语言模型(LLM)创作《哈利·波特》同人小说时,它会可靠地生成霍格沃茨宇宙的基本元素,比如可识别的地点和角色。然而,人类创作的《哈利·波特》同人小说通常既包含这些基本元素,又有更多内容,融入了风格不规则的内容和关系多样的情节线。LLM与人类写作之间的这种差距在多个领域都有被注意到:LLM倾向于生成“平均化”的写作,而人类写作包含更多样的内容,覆盖更广泛的分布。现有研究已经证实了这种分布“差距”的存在,但尚未有研究提出系统的方法来测量它。本文提出了一个以人类为基准的框架,利用人类在某一主题上写作的经验分布,来测量LLM在同一主题上生成内容的分布广度。我们提出了两个指标:LLM覆盖度(LLM-Cov)和边界内比例(IBR),将LLM内容的合理性与其分布广度分离开来。在创意构思和叙事任务中,我们发现当前的LLM生成的内容看似合理但范围狭窄,集中在人类响应空间的中心附近。我们的框架可帮助研究人员更好地评估LLM生成内容的分布广度,我们将其称为“文化覆盖范围”。

英文摘要

When a large language model (LLM) writes Harry Potter fanfiction, it reliably produces fundamental elements of the Hogwarts universe, such as recognizable places and characters. Human-written Harry Potter fanfictions, however, typically include these fundamentals and much more, incorporating stylistically irregular content and relationship-diverse plotlines. This gap between LLM and human writing has been noted across a variety of domains. LLMs tend to produce "average" writing, while human writing contains more diverse content that covers a broader distribution. Existing work has shown the existence of this distributional "gap", but no work has proposed a systematic way to measure it. Our paper proposes a human-grounded framework that uses the empirical distribution of human writing on a topic to measure the distributional breadth of LLM-generated content on that same topic. We propose two metrics, LLM Coverage (LLM-Cov) and In-Boundary Rate (IBR), that separate the plausibility of LLM content from its distributional breadth. Across ideation and narrative tasks, we find that current LLMs produce plausible but narrow content that concentrates near the center of the human response space. Our framework can enable researchers to better assess the distributional breadth of LLM-authored content, which we term its "cultural reach".

补充信息

↑