发表机构
Endgame Labs, Inc.; Asia AI Institute; Musashino University; Faculty of Data Science, Musashino University; AIx, Inc.(终局实验室公司; 亚洲人工智能研究院; 武藏野大学; 武藏野大学数据科学学院; AIx公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出摄入时语义编译(ISC)范式,将语料编译为可查询数据库对象,实验显示其检索效果优于传统RAG方法,增量更新成本更低,为RAG提供了索引化解决方案。
AI 中文摘要
几乎所有投入使用的检索增强问答系统都暗藏一个解释器:每次查询时,语言模型会重新推导原始语料文本的含义,之后便丢弃该工作。更廉价的模型并未缩小差距:每token价格已下降数个数量级,而推理支出却在上升,因为上下文规模的增长速度远超价格下降速度。这相当于现代的全表扫描,而解决办法是数据库五十年前就发现的:在写入时完成昂贵的工作,构建成可使读取变得廉价的维护结构。在语料遇到用户前若已知其读取模式,就应当对其建立索引。我们将此范式称为摄入时语义编译(ISC):将语料的含义编译为可查询的底层结构,该结构包含两个耦合层——增量维护的嵌入,以及在编译时验证来源的原子声明——并将该底层结构视为一等数据库对象,拥有自身的数据定义语言(DDL)、维护约定、迁移约定和成本模型。两项实证结果支持该范式:底层结构的维护随变化量而非语料规模扩展,增量更新的成本仅为重构的1/33.7,同时跟踪至浮点精度;在500份广播采访 transcript 的保留样本中,编译后的声明作为检索载荷,在所有32个按模型划分的预算单元中均胜出:使用约2200个读取token时准确率为85.2%,而最佳分块配置使用16300个token时准确率仅为72.5%。唯一能与之持平的基准是带有混合检索和重排序的上下文分块 pipeline,其查询路径token约为编译声明的21倍,统计上与编译声明无差异——我们认为,这种持平恰好是因为该基准自身也已开始编译。最后我们探讨由此开启的系统议程,从编译规划器到读取规划。
英文摘要
Nearly every retrieval-augmented question-answering system in production ships with a hidden interpreter: on each query a language model re-derives the meaning of raw corpus text and then throws that work away. Cheaper models do not close the gap: per-token prices have fallen by orders of magnitude while inference spend has risen, because context volume grows faster than prices fall. This is the modern equivalent of the full-table scan, and the remedy is the one databases found fifty years ago: do the expensive work once, at write time, into a maintained structure that makes reads cheap. A corpus whose read pattern is known before it ever meets a user can and should be indexed too. We call the paradigm ingest-time semantic compilation (ISC): compile a corpus's meaning into a queryable substrate with two coupled layers - incrementally maintained embeddings, and atomic claims whose provenance is validated at compile time - and treat that substrate as a first-class database object with its own DDL, maintenance contract, migration contract, and cost model. Two existence proofs support it. Substrate upkeep scales with change rather than corpus size: incremental updates run 33.7x cheaper than reconstruction while tracking it to floating-point precision. And on a held-out sample of 500 broadcast-interview transcripts, compiled claims as the retrieval payload win all 32 budget-by-model cells: 85.2% correct from roughly 2.2k reader tokens against 72.5% from 16.3k for the best chunk configuration anywhere. The only baseline that keeps pace is a contextualized-chunk pipeline with hybrid retrieval and reranking, statistically indistinguishable from compiled claims at roughly twenty-one times the query-path tokens - and it reaches that parity, we argue, precisely because it has itself begun to compile. We close with the systems agenda this opens, from compilation planners to read planning.
CommentsPosition paper. 6 pages, 2 figures, 2 tables