arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

超越模式坍缩:通过生成流网络生成多样化的合成专家对话

Beyond Mode Collapse: Generating Diverse Synthetic Expert Conversations via Generative Flow Networks

Sumit Asthana, Michael Ion, Kevyn Collins Thompson

arXiv 2609.38359首次发表:更新:

发表机构

University of Michigan(密歇根大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出用生成流网络生成多样化合成专家对话,解决数据模式坍缩问题,在辅导和情感支持领域优于强化学习和端到端LLM基线,提升下游任务训练信号。

AI 中文摘要

高质量合成数据对于后训练大型语言模型以适应型AI应用至关重要,这些应用需在对话中体现多样化的专家策略和决策。直接提示大型语言模型或将其条件化于最终使用场景,会产生低多样性数据,并坍缩至主导模式。我们提出一种使用生成流网络(GFlowNets)生成多样化高质量合成数据的方法。我们展示了训练GFlowNets以生成潜在对话结构,并使用关键交互特征(如困惑情节动态、脚手架指令平衡)上的高斯混合密度,能够按专家策略在训练数据中的流行程度比例进行采样。在结构上不同的两个领域(辅导和情感支持对话)中,与强化学习和端到端大型语言模型基线相比,我们基于GFlow的合成数据生成方法在保真度、模式覆盖和真实性方面提供了更好的平衡,且不复制训练数据。在三个下游结果预测任务上的评估表明,在合成GFlowNet生成的对话上训练的分类器比竞争性合成基线提供了更强的训练信号。

英文摘要

High quality synthetic data is central to post training LLMs for adaptive AI applications that represent the diverse expert strategies and decisions in conversations. Prompting LLMs directly or conditioning them on end use scenarios yields low diversity data that collapses onto dominant modes. We propose a method to generate diverse high quality synthetic data using Generative Flow Networks (GFlowNets). We show that training GFlowNets to generate latent conversation structure using a Gaussian mixture density over key interaction features (e.g., confusion episode dynamics, scaffolding directive balance) enables sampling expert strategies in proportion to their prevalence in the training data. Across two structurally distinct domains, tutoring and emotional support dialogues, our GFlow based synthetic data generation approach offers a better balance of fidelity, mode coverage and authenticity than reinforcement-learning and end to end LLM baselines, without copying training data. Evaluated on three downstream outcome prediction tasks, classifiers trained on synthetic GFlowNet generated conversations provide a stronger training signal than competitive synthesis baselines.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑