发表机构
Snowflake Inc.(斯诺flake公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对语义数据处理系统中LLM调用成本高、延迟大的问题,提出组合式在线学习框架,在Cortex AISQL案例中实现约8倍的实际性能提升。
AI 中文摘要
语义数据处理系统中的大语言模型(LLM)调用成本高昂,足以主导查询成本,且速度缓慢,使得CPU端学习器的更新可隐藏在其往返延迟之后。在生产环境中,LLM计算占查询成本的80%-90%,每次调用的成本是关系型谓词的10^5至10^7倍。这种延迟反转了经典自适应查询处理的设计约束,在线学习器曾需保持轻量级以避免主导其所优化的谓词;而在LLM延迟下,每次调用的梯度步骤和每批次阈值求解可容纳在往返延迟内。我们在LLM调用边界开发了组合式在线学习,这是一种用于组合语义数据处理系统中在线学习组件的框架,每个组件在执行时做出决策并在线优化其学习到的产物。设计空间涵盖决策粒度和学习器更新节奏两个维度,组件共享单一学习模式,将每个训练步骤隐藏在下一次LLM往返中。在Cortex AISQL中的生产案例研究组合了三个组件: memoization层、逐调用的在线过滤排序学习器、逐批次的在线级联路由学习器。条件成本分解将每个学习组件分配到逐行LLM成本的不同因子中;在独立情况下,两个学习组件组合后,对于代表性的合取过滤工作负载,其性能提升上限为11.4倍;级联边界的自选择、样本预算缩减及选择性估计漂移,将其降至接近8倍的实际值。
英文摘要
An LLM call in a semantic data processing system is expensive enough to dominate query cost, yet slow enough to hide a CPU-side learner's update behind its round-trip. In production, LLM compute accounts for $80-90\%$ of query cost, and each call costs $10^5-10^7\times$ a relational predicate. The latency window inverts a design constraint of classical adaptive query processing, where online learners had to stay lightweight to avoid dominating the predicates they optimize. At LLM latency, per-call gradient steps and per-batch threshold solves fit inside the round-trip. We develop compositional online learning at the LLM call boundary: a framework for combining online-learning components in semantic data processing systems. Each component makes execution-time decisions and refines its learned artifacts online. The design space spans two axes, decision granularity and learner update cadence, and the components share a single learning pattern that hides each trainer step inside the next LLM round-trip. A production case study in Cortex AISQL composes three components: a memoization layer, an online per-call filter-ordering learner, and an online per-batch cascade-routing learner. A conditional cost decomposition assigns each learning component to a distinct factor of per-row LLM cost. Under independence, the two learning components compose multiplicatively to an $11.4\times$ upper bound on a representative conjunction-filter workload. Self-selection at the cascade boundary, sample-budget shrinkage, and selectivity-estimation drift reduce it to a realistic figure near $8\times$.