arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.16621cs.AIcs.DBcs.IR

成本随变化量而非语料库规模变化:增量式维护演化中的语义基底

Cost Scales with Change, Not Corpus Size: Incrementally Maintaining an Evolving Semantic Substrate

Yusuke Takahashi, Kyle Wild, Asako Uraki

AI总结:

该研究提出语义基底维护成本随变化量而非语料库规模变化,通过增量低秩更新等方法,以远低于完整重构的成本实现语义基底的高效维护,且精度满足要求。

AI中文摘要:

检索增强与智能体问答系统日益在查询时重新推导语料库的含义。简单来说,系统并非针对每个问题重新推导语料库的含义,而是在文档到达时完成一次推导,之后仅需参考即可——这是一种意义的编译器,而非解释器。另一种方案是在摄入时将含义一次性编译为紧凑、可查询的语义基底,并在语料库演化过程中对其进行维护。核心争议在于维护成本:每次变化后重建截断奇异值分解(SVD)的成本似乎过高,且嵌入模型的变更似乎会强制要求对全部语料库重新嵌入。本文论证并通过实验证明,维护成本随变化量而非语料库规模变化。在受控合成试点中(维度256、秩32,语料库经50次更新事件从3000篇增长至9000篇),增量式低秩更新的单次更新成本比完整重新SVD低33.7倍,累计成本低23.8倍,同时增量子空间跟踪完整重新计算的精度达到浮点级别(最大主角度漂移低于1e-11度;召回@10=1.0)。正交Procrustes虚拟轴更新通过仅重新嵌入约10%的语料库,使平均余弦相似度恢复至真正重新嵌入向量的0.95。这些结果支持维护而非重复重构语义基底的思路。

英文摘要:

Retrieval-augmented and agentic question-answering systems increasingly re-derive the meaning of a corpus at query time. Put plainly, instead of re-deriving what a corpus means on every question, the work is done once when a document arrives and is thereafter merely consulted -- a compiler, not an interpreter, of meaning. An alternative is to compile that meaning once, at ingest time, into a compact, queryable semantic substrate and maintain it as the corpus evolves. The central objection is maintenance cost: rebuilding a truncated singular value decomposition (SVD) on every change appears prohibitive, and a change of embedding model seems to force a full re-embedding. We argue and show empirically that maintenance cost scales with the amount of change, not corpus size. On a controlled synthetic pilot (dimension 256, rank 32, a corpus grown from 3,000 to 9,000 documents over 50 update events), incremental low-rank updates were 33.7 times cheaper per update than full re-SVD and 23.8 times cheaper cumulatively, while the incremental subspace tracked the full recomputation to within floating-point precision (maximum principal-angle drift below 1e-11 degrees; recall@10 = 1.0). An orthogonal Procrustes virtual axis update recovered 0.95 mean cosine to truly re-embedded vectors by re-embedding only about 10 percent of the corpus. The results support maintaining, rather than repeatedly reconstructing, a semantic substrate.

补充信息

↑