发表机构
University of Science and Technology of China(中国科学技术大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出多智能体协作框架CuratorMAS,将数据集策展分解为五个可编程阶段并形成并行工作流,通过探索、检索领域知识、评估和过滤实现自动化,实验显示噪声率最高降36.03pp,F1分数最高提升8.88pp。
AI 中文摘要
高质量数据集对于可靠的机器学习至关重要,但数据集策展成本高昂且难以跨领域泛化。现有方法通常依赖人工设计的启发式规则或模型相关的信号,限制了它们在不同任务和用户查询中的适用性。为解决这些限制并实现数据策展的自动化,我们提出了CuratorMAS,一个多智能体协作框架,通过编排多个智能体来评估和策展高质量数据集。为实现灵活策展的目标,CuratorMAS将复杂的策展过程分解为五个可编程的执行阶段,并形成一个可并行化的工作流程。具体而言,CuratorMAS首先执行数据集探索,收集文件结构和约束线索等上下文信息,从而对给定任务形成全面理解。为了获取最新信息,CuratorMAS从在线来源检索领域知识以增强评估过程。接下来,CuratorMAS推导必要的评估标准并计算相应的指标。基于这些结果,CuratorMAS相应地执行过滤。最后,一个演化模块总结评估结果并更新相关技能。大量且全面的实验表明,CuratorMAS显著降低了噪声率,最高降低36.03个百分点(pp),同时将下游模型的F1分数最高提升8.88个百分点。
英文摘要
High-quality datasets are essential for reliable machine learning, but dataset curation remains costly and hard to generalize across domains. Existing methods typically rely on manually designed heuristics or model-dependent signals, limiting their applicability across tasks and user queries. To address these limitations and automate data curation, we propose \textbf{CuratorMAS}, a multi-agent collaboration framework that orchestrates agents to evaluate and curate high-quality datasets. To achieve the goal of flexible curation, CuratorMAS decomposes the complex curation process into five programmable execution stages and forms a parallelizable workflow. Specifically, CuratorMAS first performs dataset exploration to collect contextual information such as file structures and constraint cues, thereby developing a comprehensive understanding of the given task. In order to acquire up-to-date information, CuratorMAS retrieves domain knowledge from online sources to augment the evaluation process. Next, CuratorMAS derives the necessary evaluation criteria and computes the corresponding metrics. Based on these results, CuratorMAS executes filtering accordingly. Finally, an evolution module summarizes the evaluation outcomes and updates the relevant skills. Extensive and comprehensive experiments demonstrate that CuratorMAS significantly reduces the noise rate by up to 36.03 percentage points (pp) while also improving the F1 score of downstream models by up to 8.88 pp.