arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.19746cs.CL

PersonalBench:测量大语言模型个性化中的作者差距

PersonalBench: Measuring the Authorship Gap in LLM Personalization

Yash Ganpat Sawant

首次发表
浏览论文内容

中文总结 AI 辅助

本研究推出PersonalBench基准,通过LUAR、LLM评判者、自动风格计量学评估LLM个性化方法,发现其可区分作者但未达人类水平,发布该基准为LLM个性化提供校准工具。

中文摘要 AI 辅助

个性化文本生成旨在让大语言模型(LLM)以特定个人的风格写作,但现有的基准测试仅衡量任务准确率或偏好对齐,而非模型输出是否真正类似目标作者的写作风格。我们推出PersonalBench,这是一个通过三个独立维度评估推理时个性化方法的基准:LUAR(一个经过训练的作者身份验证模型)、LLM作为评判者以及自动风格计量学。在50位作者、1000个生成文本以及两个模型家族(Qwen 3、GLM-4)的实验中,我们发现个性化方法确实能产生可区分目标作者的输出(LUAR在生成文本中区分目标作者的AUC为0.918),但这种区分从未跨越人类与LLM的边界。所有方法与真实作者的LUAR相似度在0.484至0.508之间,低于人类跨作者的下限0.626(上限0.756)。LLM自身的作者身份指纹占主导地位:生成文本与任何人类作者的距离都大于随机人类之间的距离。尽管在LLM评判者上表现出差异,各方法在LUAR上的统计结果无显著差异(差值为0.024),我们将这种差异归因于特征提取与轮廓提取之间的循环性。我们验证了LUAR在我们的语料库中能可靠测量作者身份(单篇文本AUC为0.76,多篇文本AUC为0.96)。我们发布PersonalBench作为一个经过校准的测量工具:推理时个性化方法可调节LLM的风格,但无法弥合与人类作者的差距。

英文摘要

Personalized text generation aims to make LLMs write in a specific individual's style, yet existing benchmarks measure task accuracy or preference alignment rather than whether the model's output actually resembles the target author's writing. We introduce PersonalBench, a benchmark that evaluates inference-time personalization methods through three independent lenses: LUAR (a trained authorship verification model), an LLM-as-judge, and automated stylometrics. Across 50 authors, 1,000 generations, and two model families (Qwen 3, GLM-4), we find that personalization methods do produce author-differentiated output (LUAR discriminates target authors within generated text at AUC=0.918) but this differentiation never crosses the human-LLM boundary. All methods achieve LUAR similarity to real authors in the range 0.484-0.508, below the cross-author human floor of 0.626 (ceiling 0.756). The LLM's own authorship fingerprint dominates: generated text is more distant from any human author than random humans are from each other. Methods are statistically indistinguishable from each other on LUAR (spread 0.024) despite appearing differentiated on the LLM judge, a discrepancy we trace to circularity between trait extraction and profile extraction. We validate that LUAR reliably measures authorship in our corpus (AUC=0.76 single-post, 0.96 multi-post). We release PersonalBench as a calibrated measuring stick: inference-time personalization modulates the LLM's style but does not bridge the gap to human authorship.

发表机构

  • Independent AI Researcher(独立人工智能研究员)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑