DataClaw0: 从原始流中智能定制多模态数据
DataClaw0: Agentic Tailoring Multimodal Data from Raw Streams
- University of Chinese Academy of Sciences(中国科学院大学)
- Shenzhen University of Advanced Technology(深圳理工大学)
- Tsinghua University(清华大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
提出DataClaw0模型,通过两阶段流水线将生成语义合成锚定于确定性事实锚点,结合SFT与GRPO实现复杂数据精炼意图的对齐,首个数据精炼基准验证其高效性。
AI中文摘要:
大规模非结构化多模态流存在高“数据熵”,阻碍了高效的人类知识获取和高质量的AI后训练。现有的被动标注范式严重依赖启发式规则或通用VLM,成本高、单调且无法解锁原始数据中嵌入的深层过程逻辑。我们将数据处理提升为可学习的能力,提出向智能数据定制的范式转变,主动精炼和结构化数据以对齐多样化的用户和下游意图。为了克服训练这种高阶能力的数据稀缺瓶颈,我们设计了一个两阶段流水线,将生成语义合成锚定于确定性事实锚点,生成覆盖五个核心物理和数字领域的大规模数据集。在此基础上,$\text{DataClaw}_0$-9B模型协同监督微调(SFT)与组相对策略优化(GRPO),实现了与复杂精炼和定制意图的稳健对齐。为了系统量化这一能力,我们构建了$\text{DataClaw}_0$-val,这是首个专门用于数据精炼的基准。关键的是,我们采用下游后训练作为最终的验证试金石。在视频生成、真实世界VQA和GUI导航上的评估证实,$\text{DataClaw}_0$提供高信息密度的定制数据,促进模型在有限训练数据下高效适应新任务。项目页面:this https URL
英文摘要:
Raw multimodal streams are abundant but noisy, redundant, and unaligned with any particular training objective. Turning them into supervision today means either brittle heuristics or repeatedly querying a proprietary vision-language model, a cost that recurs with every new sample. We ask whether this conversion can instead be learned once and reused, and formalise intent-conditioned Data Tailoring: given a raw stream and a high-level intent, a model must return schema-aligned, evidence-grounded training instances. Training DataClaw0 at 4B, 9B and 27B, we find that whether five heterogeneous domains should share one model depends on capacity. A jointly trained model is worse than per-domain experts at the two smaller scales and better at the largest, placing the crossover near 18B parameters. Matched-data comparisons, in which the joint model sees exactly the same data per domain as that domain's expert, attribute the reversal to cross-domain transfer rather than to data volume: it gains most where a domain is data-poor, and recovers 59\% of in-domain performance on domains withheld from training entirely. Downstream post-training reproduces this ordering on GUI navigation, action video generation and spatio-temporal VQA, and the joint configuration is also the cheaper to deploy, serving one model instead of five. Github: https://github.com/vancyland/DataClaw0