AI 中文总结
提出SemBaker编译式执行引擎,通过生成Python函数本地执行替代逐项LLM调用,支持多工具,在200查询的问答工作负载中获4.8-6.3倍加速与5.4-10.7倍成本降低。
AI 中文摘要
语义算子通过自然语言谓词扩展了数据处理能力。现有语义算子系统通常采用基于解释的执行方式来运行这些算子:对于每个数据项,大型语言模型(LLM)会解释算子谓词并直接生成对应结果。尽管这种方式表达能力强,但它将昂贵的模型调用置于数据处理循环中,导致延迟和资金成本随输入基数成比例增加。我们提出SemBaker,一种面向语义算子系统的基于编译的执行引擎。SemBaker作为外部插件而非替换后端的原生执行,针对选定的语义过滤、映射和连接操作,它仅调用一次LLM生成确定性Python函数,随后在本地执行该函数,无需对每个数据项调用LLM。基于成本的优化器会将每个算子路由至原生或编译执行,同时编译过程与流水线执行重叠。SemBaker通过轻量适配器支持Palimpzest、LOTUS、Nirvana和DocETL。在三个包含200个查询的问答工作负载中,SemBaker实现了4.8至6.3倍的平均加速比和5.4至10.7倍的平均成本降低,且处理质量具有竞争力。
英文摘要
Semantic operators extend data processing with natural-language predicates. Existing semantic operator systems commonly execute these operators through interpretation-based execution: for every data item, an LLM interprets the operator predicate and directly produces the corresponding result. Although expressive, this design places expensive model invocations inside the data-processing loop, causing latency and monetary cost to scale with input cardinality. We present SemBaker, a compilation-based execution engine for semantic operator systems. SemBaker acts as an external plugin rather than replacing a backend's native execution. For selected semantic filters, maps, and joins, it invokes an LLM once to generate a deterministic Python function and executes that function locally without per-item LLM calls. A cost-based optimizer routes each operator to native or compiled execution, while compilation overlaps pipeline execution. SemBaker supports Palimpzest, LOTUS, Nirvana, and DocETL through thin adapters. Across three 200-query QA workloads, SemBaker achieves average speedups of 4.8 to 6.3 times and average cost reductions of 5.4 to 10.7 times, with competitive processing quality.