arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

边际效用、矩阵分解与键值(KV)缓存:面向主权地理采矿推理的统一信息经济学框架

Marginal utility, matrix factorization, and the Key-Value (KV) cache: a unified information-economic framework for sovereign geo-mining inference

Caroline Gans Combe

arXiv 2609.20068首次发表:更新:

发表机构

INSEEC Business School / Omnes Education(INSEEC商学院 / Omnes教育集团)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出统一信息经济学框架,将边际效用与矩阵分解及KV缓存结合,用于主权地理采矿推理,并验证了高效分类器与合并策略。

AI 中文摘要

本文在经济学中的边际效用概念与两种机器学习构造——矩阵分解和Transformer语言模型的键值(Key-Value, KV)缓存——之间建立了理论桥梁。评分矩阵的奇异值谱被证明是潜在因子的边际效用递减表,投影协方差算子的特征值谱是模型学习表示的边际效用表,而缓存逐出和低秩缓存压缩则是内存预算约束下的约束效用最大化实例。这三者归结为单一的分配规则:保留特征值超过约束影子价格的顶部维度。该框架被应用于从地理采矿文档中自动提取结构化信息,在此背景下,它催生了一种多遍推理协议、一种逐层TIES模型合并过程,以及一种结合提取质量、定位漂移和能量、并通过漂移上的条件风险价值(Conditional Value-at-Risk, CVaR)项进行标量化的选择策略。报告了两项实证贡献。一个1120万参数的层次分类器,在单个GPU上约五分钟内训练完成,在来自973篇铀勘探语料库的留出测试集上达到了90.0%的一级准确率,而一个专有模型在同一语料库的五十篇文档人工审计中达到92.0%,该分类器每张卡片的延迟为2.62毫秒,而API约为2000毫秒,且成本可忽略不计。对均匀密度TIES合并的诊断揭示了一种可复现的退化模式,在该模式下,合并后的模型在五个地理上不同的区域返回令牌相同的输出,同时声明高置信度;在逐层校准密度下重新执行合并,消除了诊断样本上的该特征。全规模提取基准(包括LoRA微调)被报告为预计值而非实测值,并且仍然是本工作的实证扩展。

英文摘要

This paper builds a theoretical bridge between the economic notion of marginal utility and two machine-learning constructs, matrix factorization and the Key--Value cache of transformer language models. The singular value spectrum of a rating matrix is shown to be a diminishing marginal utility schedule for latent factors, the eigenvalue spectrum of the projected covariance operator to be the marginal utility schedule of a model's learned representation, and cache eviction and low-rank cache compression to be instances of constrained utility maximization under a memory budget. The three collapse into a single allocation rule: retain the top dimensions whose eigenvalue exceeds the shadow price of the binding constraint. The framework is applied to the automated extraction of structured information from geo-mining documents, where it motivates a multi-pass inference protocol, a layer-wise TIES model merging procedure, and a selection policy combining extraction quality, localization drift and energy, scalarized with a Conditional Value-at-Risk term on drift. Two empirical contributions are reported. An 11.2-million-parameter hierarchical classifier, trained in about five minutes on a single GPU, reaches 90.0 per cent level-1 accuracy on a held-out test set from a 973-document uranium-exploration corpus, against 92.0 per cent for a proprietary model on a fifty-document human audit of the same corpus, at a latency of 2.62 ms per card against approximately 2,000 ms for the API and at negligible cost. A diagnostic of uniform-density TIES merging exposes a reproducible degenerate mode in which the merged model returns token-identical outputs across five geographically distinct districts while declaring high confidence; re-executing the merge under layer-wise calibrated densities removes that signature on the diagnostic sample. The full-scale extraction benchmark, including LoRA fine-tuning, is reported as projected rather than measured and remains an empirical extension of this work.

CommentsVersion 11, 14 septembre 2026. 49 pages, 9 tables. Les valeurs de l'architecture souveraine sont projet{é}es et non mesur{é}es ; le calcul {à} grande {é}chelle est en cours. Soumission pr{é}vue {à} IEEE Transactions on Artificial Intelligence

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑