arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.03652cs.CL

合成数据增强对话-语用功能分类的影响

The Impact of Synthetic Data Augmentation on Discourse-Pragmatic Function Classification

Sara Sorahi, Kevin Tang, Reza Kazemian

首次发表
浏览论文内容

中文总结 AI 辅助

该研究在对话语用功能分类任务中,以Llama 3.1生成合成样本并按RoBERTa嵌入空间的余弦距离划分,发现核心近端样本提升宏观F值最多,距离平衡混合样本提升准确率最高,表明合成样本的表示空间位置对分类效果影响显著。

中文摘要 AI 辅助

合成数据增强已成为解决自然语言处理(NLP)中类别不平衡问题的常用策略,但大多数方法关注生成样本的数量和多样性,而非其与真实训练数据的几何关系。我们在对话语用功能分类任务中研究这一问题,该任务的数据稀疏性是结构特征而非采集人工制品。使用来自英国国家语料库的410个手动标注的英语单词look实例,涵盖注意力信号、指令、话语标记语和感叹词四种功能。我们使用Llama 3.1生成合成训练样本,并按其在RoBERTa嵌入空间中与真实训练数据的余弦距离划分。我们比较六种训练条件,这些条件在合成样本相对于经验决策边界的位置上存在差异,同时保持各条件间的增强数量恒定。所有增强条件均比仅使用真实数据的基线提高了宏观F值和准确率,但核心近端样本(NEAR)带来的宏观F值提升最大(0.113),而距离平衡混合样本实现了最高的准确率(0.748)。没有任何条件提高AUC,表明增强会改变决策边界,而非改善模型的基础概率估计。这些发现表明,合成样本在表示空间中的位置与生成数量同样重要,对更广泛的低资源语用分类具有启示意义。

英文摘要

Synthetic data augmentation has become a common strategy for addressing class imbalance in NLP, but most approaches focus on the quantity and diversity of generated examples rather than their geometric relationship to real training data. We investigate this question in the context of discourse pragmatic function classification, a task where data sparsity is a structural feature rather than a collection artefact. Using 410 manually annotated instances of the English word look drawn from the British National Corpus, spanning four functions: Attention Signal, Directive, Discourse Marker, and Interjection. We generate synthetic training examples with Llama 3.1 and partition them by their cosine distance from real training data in RoBERTa embedding space. We compare six training conditions that differ in the placement of synthetic examples relative to the empirical decision boundary, while holding augmentation quantity constant across conditions. All augmented conditions improve macro F and accuracy over the real only baseline, but core proximal examples (NEAR) yield the largest gains in macro F (0.113), while a distance balanced mix achieves the highest accuracy (0.748). No condition improves AUC, indicating that augmentation shifts the decision boundary rather than improving the model's underlying probability estimates. These findings suggest that where synthetic examples land in representation space matters as much as how many are generated, with implications for low resource pragmatic classification more broadly.

发表机构

  • Institute of Linguistics(语言学研究所)
  • College of Liberal Arts and Sciences, University of Florida(佛罗里达大学文理学院)
  • Faculty of Arts and Humanities, Heinrich Heine University Düsseldorf(杜塞尔多夫海因里希·海涅大学艺术与人文学院)
  • Sun Yat-sen University(中山大学)

机构由 AI 辅助整理,请以论文原文为准。

↑