arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

隐私个性化在LLMs中的权衡:风格测量信号减少对用户特定文本生成的影响

Privacy Personalization Trade offs in LLMs: The Impact of Stylometric Signal Reduction on User-Specific Text Generation

Muhammed Nazmul Arefin, Omar Jamal Hammad

arXiv 2609.22112首次发表:更新:

发表机构

King Fahd University of Petroleum and Minerals(法赫德国王石油与矿业大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过LaMP-7基准实验,发现减少风格信号会显著降低LLM个性化文本的偏好(降至13.0%),但语义保留高(94.8%),揭示了隐私与个性化之间的权衡。

AI 中文摘要

大型语言模型(LLMs)已展现出以高风格保真度生成用户特定文本的能力。然而,实现这种个性化的个人数据常常嵌入人口统计、文化和风格标记,引发了对风格测量重新识别的担忧。本文研究减少可识别的风格信号是否影响LLMs在文本生成中的个性化。我们引入了一个受控框架,使用LaMP-7 Twitter基准来隔离LLM个性化中的风格测量信号。对250个采样用户的实验比较了两种设置:基于原始档案的条件改写和基于匿名化转换档案的条件改写,后者中人口统计标识符、文化参考、个人细节和非正式语言线索已被系统性地中和。输出由两个独立的LLM评审员和一项补充性人工评估进行评估。我们的成对评估显示,基于原始档案的输出与人类撰写的真实文本几乎无法区分,表明现代LLMs能够以足够的保真度紧密重现作者的写作风格。相比之下,使用匿名化档案的模型输出的偏好平均降至13.0%,而语义上下文保留仍高达94.8%。一项由人工评估者进行的研究证实了同样的模式。这些发现揭示了明确的隐私-个性化权衡,并强调需要开发隐私感知的个性化方法,在保留意义的同时抑制识别性风格信号。

英文摘要

Large language models (LLMs) have demonstrated the ability to generate user-specific text with high stylistic fidelity. However, the personal data that enables such personalization frequently embeds demographic, cultural, and stylistic markers that raises concerns about stylometric re- identification. This paper investigates whether reducing identifiable stylistic signals affects personalization in text generation by LLMs. We introduce a controlled framework to isolate stylometric signals in LLM personalization using the LaMP-7 Twitter benchmark. Experiments on 250 sampled users compare two settings: paraphrasing conditioned on the original profile and paraphrasing conditioned on an anonymized converted profile in which demographic identifiers, cultural references, personal details, and informal linguistic cues have been systematically neutralized. Outputs are assessed by two independent LLM judges and a complementary human evaluation. Our pairwise evaluation shows that outputs conditioned on original profiles are nearly indistinguishable from human-authored ground truth, indicating that modern LLMs can closely reproduce an author's writing style with sufficient fidelity. In contrast, preference for model outputs with anonymized profiles drops to 13.0% on average, while semantic context preservation remains high at 94.8%. A study with human evaluators confirms the same pattern. These findings reveal a clear privacy-personalization trade-off and highlight the need for privacy-aware personalization methods that retain meaning while suppressing identifying stylistic signals.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑