发表机构
East China Normal University; The Hong Kong University of Science and Technology (Guangzhou); Shanghai Innovation Institute(华东师范大学; 香港科技大学(广州); 上海创新研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对大语言模型与以人为中心目标对齐的难题,提出以数据为中心的ADE框架,通过OVS流程与稳态准入机制提升数据质量,在多基准及不同场景下均实现性能提升。
AI 中文摘要
当目标不可执行且依赖上下文时,将大语言模型与以人为中心的目标对齐十分困难,这限制了可靠验证和可扩展监督。尽管合成数据扩大了覆盖范围,但弱验证使瓶颈从生成转向选择,噪声信号破坏迭代优化并可能导致隐性退化。我们提出Agentic Data Evolution(ADE,智能体数据演化),一种以数据为中心的框架,将合成监督组织为演化数据快照。ADE通过闭环观察-变异-选择(OVS)流程改进数据快照,其中稳态准入机制作为质量棘轮,保守地管控更新以实现持续跨轮改进。我们通过互补的内在趋势跟踪和外在后训练评估验证这些改进:在DEV300上,ADE将内在胜率从50%提升至75.81%,外在胜率从55.20%提升至68.86%,在不同基准上均表现出一致的性能提升;盲专家评估进一步证实,演化答案获得66.11%的偏好。这些提升在多种后训练方法、模型规模及超出目标弱可验证教育目标的任务中均成立。资源可在指定URL获取。
英文摘要
Aligning large language models to human-centered objectives is difficult when targets are non-executable and context-dependent, limiting reliable verification and scalable supervision. Although synthetic data expands coverage, weak verification shifts the bottleneck from generation to selection. Noisy signals destabilize iterative refinement and can cause silent regressions. We propose Agentic Data Evolution (ADE), a data-centric framework that organizes synthetic supervision as evolving data snapshots. ADE improves data snapshots through a closed-loop Observation-Variation-Selection (OVS) procedure, where a steady-state admission mechanism acts as a quality ratchet that conservatively gates updates for sustained cross-round improvement. We validate these improvements through complementary intrinsic trend tracking and extrinsic post-training evaluation. On DEV300, ADE raises the intrinsic win rate from 50% to 75.81% and the extrinsic win rate from 55.20% to 68.86%, consistent performance gains across diverse benchmarks. Blind expert evaluation further confirms this, with a 66.11% preference for evolved answers. These gains extend across post-training methods, model scales, and tasks beyond the target weakly verifiable educational objectives. Resources are available at https://github.com/ZeroLoss-Lab/Agentic-Data-Evolution.
Commentsaccepted by EMNLP 2026