一个模型不是群体:多LLM与方面条件化的多样化评论生成
One Model Is Not a Crowd: Multi-LLM and Aspect-Conditioned Diverse Comment Generation
浏览论文内容
中文总结 AI 辅助
本研究提出结合多个LLM与方面条件化生成来提升评论多样性,通过大规模YouTube评论分析,证明该方法更贴近人类分布且对下游任务有效,但人类多样性仍难以完全复现。
中文摘要 AI 辅助
互联网上的人类交流受到多元视角的影响,这在在线评论空间中表现得最为明显。随着基于大型语言模型(LLM)的AI智能体开始进入这些空间,一个关键问题浮现:合成评论线程能否捕捉人类话语中固有的多样性。这一担忧日益重要,因为同质化AI生成内容的增多可能随时间减少多样性,潜在地导致模型崩溃并削弱数字交流的丰富性。受人类群体多元性和话语方面驱动特性的启发,我们假设评论多样性通过结合多个LLM与方面条件化生成能得到更好的近似。我们使用来自不同提供商的模型对此方法进行形式化并评估,引入了一个框架,该框架在语义、语言和社会语用特征上,沿三个轴——离散度、覆盖度和对齐度——表征多样性。利用此框架,我们对跨多个领域的超过200万条YouTube评论进行了大规模研究。我们的结果显示,多LLM和方面条件化生成更好地与人类评论分布对齐,且此类数据在预训练风格的筛选下仍然可行,并对下游任务有效。然而,人类的多样性仍无法匹敌。总体而言,我们的发现为在AI中介环境中生成更多样化和具有社会基础的话语提供了实用基础。
英文摘要
Human communication on the internet is shaped by diverse perspectives, most visibly expressed in online comment spaces. As large language model (LLM)based AI agents begin to inhabit these spaces, a key question arises: whether synthetic comment threads can capture the diversity inherent in human discourse. This concern is increasingly important, as the growing presence of homogenized AI-generated content risks reducing diversity over time, potentially leading to model collapse and degrading the richness of digital communication. Inspired by the plurality of human crowds and the aspect-driven nature of discourse, we hypothesize that comment diversity is better approximated by combining multiple LLMs with aspect-conditioned generation. We formalize and evaluate this approach using models from different providers and introduce a framework that characterizes diversity across semantic, linguistic, and socio-pragmatic features along three axes: dispersion, coverage, and alignment. Using this framework, we conduct a large-scale study on over 2 million YouTube comments across multiple domains. Our results reveal that multi-LLM and aspect-conditioned generation better align with human comment distributions and such data remains viable under pretraining style curation and is effective for downstream tasks. Yet, human diversity remains unmatched. Overall, our findings provide a practical foundation for generating more diverse and socially grounded discourse in AI-mediated environments.
发表机构
- Pennsylvania State University(宾夕法尼亚州立大学)
- University of Sheffield(谢菲尔德大学)
机构由 AI 辅助整理,请以论文原文为准。