发表机构
AntGroup(蚂蚁集团)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
OmniTable提出统一宽表架构,通过逻辑统一与物理分离、声明式特征管理和自适应执行引擎,实现PB级LLM数据整理与探索,将整理周期缩短5.6倍。
AI 中文摘要
数据整理是工业级大语言模型(LLM)开发中的关键瓶颈,其中PB级非结构化语料分散在数百张物理表中,特征工程依赖人工的、以表为中心的流水线编排,且数据血缘基本缺失。我们提出OmniTable,作为一个基于逻辑统一、物理分离的统一宽表层架构蓝图,面向PB级LLM数据整理与探索。OmniTable做出四项贡献:(1)统一的宽表抽象,通过逻辑-物理映射,将多源异构数据和数千个派生特征整合到单一逻辑模式之下;(2)声明式特征生命周期管理,自动化依赖解析、执行规划、算子融合和血缘追踪,以“声明-执行”范式取代人工流水线编排;(3)具备自主治理的自适应执行引擎,通过异构计算路由(CPU/GPU)、自适应调优、UDF级容错和自动化存储布局优化,实现稳定的PB级特征回填;(4)混合加速的数据探索,结合全局ID索引、透明OLAP卸载和后台物化视图,实现秒级点查和超过20 TB/小时的过滤导出。在生产环境中,OmniTable管理超过35 PB的训练数据,涵盖网页、代码、PDF和SFT领域,将人在回路的整理周期从约14天缩短至约2.5天(较OmniTable之前的生产工作流提升5.6倍),并保持一致的特性版本、可审计的血缘和最少的人工干预。
英文摘要
Data curation is a critical bottleneck in industrial-grade LLM development, where petabyte-scale unstructured corpora are scattered across hundreds of physical tables, feature engineering relies on manual, table-centric pipeline orchestration, and data lineage is largely absent. We present OmniTable as an architecture blueprint for a unified wide-table layer built on Logical Unification, Physical Separation, targeting petabyte-scale LLM data curation and exploration. OmniTable makes four contributions: (1) a unified wide-table abstraction that consolidates multi-source heterogeneous data and thousands of derived features under a single logical schema via logical-physical mapping; (2) declarative feature lifecycle management that automates dependency resolution, execution planning, operator fusion, and lineage tracking, replacing manual pipeline orchestration with a "declare-and-execute" paradigm; (3) an adaptive execution engine with autonomous governance that achieves stable PB-scale feature backfill through heterogeneous compute routing (CPU/GPU), adaptive tuning, UDF-level fault tolerance, and automated storage layout optimization; and (4) hybrid-accelerated data exploration combining a global ID index, transparent OLAP offloading, and background materialized views to deliver second-level point lookups and filtered exports exceeding 20 TB/hour. In production, OmniTable manages over 35 PB of training data across web, code, PDF, and SFT domains, reducing the human-in-the-loop curation cycle from approximately 14 days to approximately 2.5 days (5.6x over the pre-OmniTable production workflow), with consistent feature versioning, auditable lineage, and minimal manual intervention.
CommentsVLDB 2026 Best Industry Paper
Journal refProceedings of the VLDB Endowment 19(12):4276-4289, 2026