arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.17549cs.CLcs.CY

社会模式在合成数据中是否成立?分析LLM生成对话与真实对话中的网络欺凌动态

Do Social Patterns Hold in Synthetic Data? Analyzing Cyberbullying Dynamics in LLM-Generated and Authentic Dialogues

Arefeh Kazemi, Hamza Qadeer, Sinan Asci, Joachim Wagner, Brian Davis

首次发表
浏览论文内容

中文总结 AI 辅助

本研究评估LLM生成网络欺凌对话的社会真实性,发现其保留全局互动结构但扭曲细粒度社会动态,且扭曲程度因模型而异,合成数据在行为真实性上仍不完美。

中文摘要 AI 辅助

网络欺凌(CB)是一种复杂的社会现象,其特征为重复性攻击、权力不平衡和多主体互动。尽管大型语言模型(LLMs)越来越多地被用于生成合成网络欺凌对话,以进行数据增强和基准测试,但此类数据是否在支持下游任务性能之外,忠实地再现了真实互动的社会动态,仍不清楚。我们提出了一个全面的框架,用于评估LLM生成的网络欺凌对话的社会真实性。我们比较了由GPT、Grok和LLaMA生成的真实对话与合成对话,涵盖互动结构(话轮转换、权力动态和修复行为)、语言和风格真实性(代词使用和幽默)、情感和行为标记(网络欺凌类型、脏话和毒性)以及时间升级动态。我们进一步将自动分析与人工评估相结合,评估网络欺凌的存在、场景相关性、角色合理性和社会真实性。我们的结果表明,LLM生成的数据始终保留了高层次的互动结构,包括角色参与模式、方向性权力不对称以及行为标记的广泛分布。然而,所有模型都系统地扭曲了更细粒度的社会现象,包括行为强度、角色特定分配、类别分布和时间动态。这些扭曲强烈依赖于模型:GPT抑制有害内容,Grok放大攻击性行为,而LLaMA在平滑角色区别的同时提供了最平衡的近似。我们的发现表明,合成网络欺凌数据对于建模全局互动结构是有用的,但在行为真实性和社会动态至关重要的场景中,它仍然是真实对话的不完美替代品。

英文摘要

Cyberbullying (CB) is a complex social phenomenon characterized by repeated aggression, power imbalance, and multi-party interaction. Although large language models (LLMs) are increasingly used to generate synthetic CB conversations for data augmentation and benchmarking, it remains unclear whether such data faithfully reproduces the social dynamics of authentic interactions beyond supporting downstream task performance. We present a comprehensive framework for evaluating the social realism of LLM-generated CB conversations. We compare authentic and synthetic dialogues generated by GPT, Grok, and LLaMA across interactional structure (turn-taking, power dynamics, and repair behavior), linguistic and stylistic realism (pronoun usage and humor), affective and behavioral markers (CB types, profanity, and toxicity), and temporal escalation dynamics. We further complement automatic analyses with a human evaluation of cyberbullying presence, scenario relevance, role plausibility, and social realism. Our results show that LLM-generated data consistently preserves high-level interactional structure, including role participation patterns, directional power asymmetry, and broad distributions of behavioral markers. However, all models systematically distort finer-grained social phenomena, including behavioral magnitude, role-specific allocation, categorical distributions, and temporal dynamics. These distortions are strongly model-dependent: GPT suppresses harmful content, Grok amplifies aggressive behaviors, and LLaMA provides the most balanced approximation while smoothing role distinctions. Our findings show that synthetic CB data is useful for modeling global interactional structure but remains an imperfect substitute for authentic conversations when behavioral realism and social dynamics are essential.

发表机构

  • Dublin City University(都柏林城市大学)
  • DCU Anti Bullying Centre(都柏林城市大学反欺凌中心)

机构由 AI 辅助整理,请以论文原文为准。

↑