AI 中文总结
研究语义算子系统执行效率问题,提出基于编译的执行方法,在编译时调用LLM将算子规范转为可执行代码,本地运行减少调用次数,应用于多种算子并集成到系统,初步结果显示减少了执行时间和调用次数,还拓展了LLM在该领域的角色认知。
AI 中文摘要
语义算子系统通过自然语言接口扩展数据处理,支持语义过滤、映射和连接等操作。现有系统通常通过基于解释的执行来运行这些算子,这导致在数据处理循环中进行昂贵的大语言模型(LLM)调用,造成高延迟、高成本和有限的可扩展性。我们提出基于编译的语义算子执行方法,在编译期间调用LLM一次,将语义算子规范转换为确定性可执行代码。生成的代码作为编译后的物理算子,在数据集上本地运行,无需逐行或逐对调用LLM。我们将此方法应用于语义过滤、映射和连接算子,并与基于LLM的解释执行进行比较,还将其集成到现有语义算子系统中。初步结果表明,基于编译的执行大幅减少了执行时间和LLM调用,同时保留了大部分输出质量。更广泛地说,我们认为语义算子系统应将LLM不仅视为运行时执行器,还应视为生成高效可执行计划的语义编译器,这在数据库查询处理、程序合成和LLM驱动的数据系统的交叉领域开辟了新的研究机会。
英文摘要
Semantic operator systems extend data processing with natural-language interfaces, supporting operations such as semantic filtering, mapping, and joining. Existing systems commonly execute these operators through interpretation-based execution: for each row, record, or candidate pair, an LLM is invoked to interpret the semantic intent and produce an output. Although expressive, this places expensive LLM calls inside the data-processing loop, causing high latency, monetary cost, and limited scalability. We propose compilation-based execution of semantic operators. Instead of using an LLM as a runtime interpreter for every data item, we invoke it once during compilation to translate a semantic operator specification into deterministic executable code. The generated code serves as a compiled physical operator that approximates the behavior of the original LLM-based operator and runs locally over the dataset without per-row or per-pair LLM calls. We instantiate this approach for semantic filter, semantic map, and semantic join operators, compare it with LLM-based interpreted execution, and integrate it into an existing semantic operator system. Preliminary results show that compilation-based execution substantially reduces execution time and LLM calls while preserving much of the output quality. More broadly, we argue that semantic operator systems should treat LLMs not only as runtime executors, but also as semantic compilers that generate efficient executable plans, opening new research opportunities at the intersection of database query processing, program synthesis, and LLM-powered data systems.