arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.29144cs.CL

量化合成数据的误差容忍度:原子级操作数与算子扰动研究

Quantifying Error Tolerance in Synthetic Data: An Atomic-level Operand vs. Operator Perturbation Study

  • CDL, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所CDL)
  • School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院)
  • Fujian Provincial Cancer Hospital(福建省肿瘤医院)
  • College of Computer and Data Science, Fuzhou University(福州大学计算机与数据科学学院)

机构由 AI 辅助整理,请以论文原文为准。

Jiaxiang Liu, Chenhao Yuan, Shuwen Xu, Boxuan Xing, Xiusheng Huang, Yinhao Xu, Hao Liu, Wenhao Teng, Xiangwen Liao, Pengfei Cao, Jun Zhao, Kang Liu

中文总结 AI 辅助

针对合成数据误差容忍度定量分析缺失的问题,提出ATOM框架区分操作数与算子扰动,发现模型对操作数扰动鲁棒但受算子扰动影响,合成数据较LIMA提升3.1%,凸显算子多样性的重要性。

中文摘要 AI 辅助

合成数据生成已成为推动大型语言模型发展的核心支柱,但对误差容忍度的定量分析缺失成为关键瓶颈,导致当前过滤策略走向两个极端:要么过于激进,存在排除潜在有价值样本的风险;要么过于宽松,无法有效消除错误样本。为填补这一空白,本文提出原子树操作建模(ATOM)框架,该框架将数据分解为功能单元(f(x)→y),区分良性操作数x扰动与致命算子f扰动:前者会被激进过滤不必要地丢弃,后者则会被宽松过滤遗漏。实验揭示了双重解离:模型对操作数扰动具有鲁棒性,但在算子扰动下会崩溃。通过优先保证算子精度而非激进的操作数精度,经ATOM合成的数据优于严格基线(例如,较LIMA提升3.1%),表明算子多样性比操作数精度更重要。我们的代码可在该https URL获取。

英文摘要

Synthetic data generation has become a cornerstone for advancing large language models. However, the lack of the quantitative analysis for error tolerance became a critical bottleneck. Consequently, current filtering strategies fluctuate between two extremes: they are either overly aggressive, risking the exclusion of potentially valuable samples, or overly permissive, failing to eliminate erroneous samples effectively. To bridge this gap, this paper introduces Atomic Tree Operation Modeling (ATOM), a framework that decomposes data into functional units ($f(x)\rightarrow y$). ATOM distinguishes benign Operand $x$ perturbations from fatal Operator $f$ perturbations. The former are needlessly discarded by aggressive filtering, while the latter slip through permissive filtering. Our experiments reveal a double dissociation: models are robust to operand perturbations but collapse under operator perturbations. By prioritizing operator over aggressive operand precision, our ATOM-synthesized data outperforms rigorous baselines (e.g., +3.1% gain over LIMA), suggesting that operator diversity matters more than operand precision. Our code is available at https://github.com/Lut-hub/ATOM.

↑