arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

OmniTable:面向PB级LLM数据整理与探索的统一宽表系统

OmniTable: A Unified Wide-Table System for Petabyte-Scale LLM Data Curation and Exploration

Yuzhuo Fu, Xiangchun Wang, Chao Huang, Liyi Wang, Binwei Zeng, Yuhan Wang, Taotao Nie, Dongke Hu, Wang Hong, Jiayi Wang, Wenwen Cui, Zhuyan Zhou, Yushun Guo, Yuhan Xing, Jiaxin Lian, Peng Lin, Qing Cui, Wenhui Shi, Jun Zhou

arXiv 2609.11148首次发表:更新:

发表机构

AntGroup(蚂蚁集团)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

OmniTable提出统一宽表架构,通过逻辑统一与物理分离、声明式特征管理和自适应执行引擎,实现PB级LLM数据整理与探索,将整理周期缩短5.6倍。

AI 中文摘要

数据整理是工业级大语言模型(LLM)开发中的关键瓶颈,其中PB级非结构化语料分散在数百张物理表中,特征工程依赖人工的、以表为中心的流水线编排,且数据血缘基本缺失。我们提出OmniTable,作为一个基于逻辑统一、物理分离的统一宽表层架构蓝图,面向PB级LLM数据整理与探索。OmniTable做出四项贡献:(1)统一的宽表抽象,通过逻辑-物理映射,将多源异构数据和数千个派生特征整合到单一逻辑模式之下;(2)声明式特征生命周期管理,自动化依赖解析、执行规划、算子融合和血缘追踪,以“声明-执行”范式取代人工流水线编排;(3)具备自主治理的自适应执行引擎,通过异构计算路由(CPU/GPU)、自适应调优、UDF级容错和自动化存储布局优化,实现稳定的PB级特征回填;(4)混合加速的数据探索,结合全局ID索引、透明OLAP卸载和后台物化视图,实现秒级点查和超过20 TB/小时的过滤导出。在生产环境中,OmniTable管理超过35 PB的训练数据,涵盖网页、代码、PDF和SFT领域,将人在回路的整理周期从约14天缩短至约2.5天(较OmniTable之前的生产工作流提升5.6倍),并保持一致的特性版本、可审计的血缘和最少的人工干预。

英文摘要

Data curation is a critical bottleneck in industrial-grade LLM development, where petabyte-scale unstructured corpora are scattered across hundreds of physical tables, feature engineering relies on manual, table-centric pipeline orchestration, and data lineage is largely absent. We present OmniTable as an architecture blueprint for a unified wide-table layer built on Logical Unification, Physical Separation, targeting petabyte-scale LLM data curation and exploration. OmniTable makes four contributions: (1) a unified wide-table abstraction that consolidates multi-source heterogeneous data and thousands of derived features under a single logical schema via logical-physical mapping; (2) declarative feature lifecycle management that automates dependency resolution, execution planning, operator fusion, and lineage tracking, replacing manual pipeline orchestration with a "declare-and-execute" paradigm; (3) an adaptive execution engine with autonomous governance that achieves stable PB-scale feature backfill through heterogeneous compute routing (CPU/GPU), adaptive tuning, UDF-level fault tolerance, and automated storage layout optimization; and (4) hybrid-accelerated data exploration combining a global ID index, transparent OLAP offloading, and background materialized views to deliver second-level point lookups and filtered exports exceeding 20 TB/hour. In production, OmniTable manages over 35 PB of training data across web, code, PDF, and SFT domains, reducing the human-in-the-loop curation cycle from approximately 14 days to approximately 2.5 days (5.6x over the pre-OmniTable production workflow), with consistent feature versioning, auditable lineage, and minimal manual intervention.

CommentsVLDB 2026 Best Industry Paper

Journal refProceedings of the VLDB Endowment 19(12):4276-4289, 2026

DOI:10.14778/3827998.3828032

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑