arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.05863cs.AI

异构注意力内存的运行时可观测性

Runtime Observability for Heterogeneous Attention Memory

发表机构Metask实验室
查看机构详情
  • Metask Lab(Metask实验室)

机构由 AI 辅助整理,请以论文原文为准。

Fanzhe Wei, Li Liu, Ziyang Wang, Chenyu Wang

首次发表
浏览论文内容

中文总结 AI 辅助

本研究针对异构注意力内存提出运行时可观测性契约,在多模型配置上实例化并构建风险账本,成功定位DeepSeek-V4栈的静默故障,相关成果及工具已公开。

中文摘要 AI 辅助

现代模型不再保留普通的KV缓存:潜在缓存、学习到的稀疏选择器和循环状态分别以不同形式承载模型的内存,且每种在压缩下的失效方式各不相同。我们提出了一种运行时可观测性契约,该契约涵盖四类内存并包含三个算子,将其实例化到五个架构族的六种模型配置上,并将各阶段边界组合成可执行的请求级风险账本。契约将其误差度量作为一种类型——仅当度量匹配时才定义组合,此检查拒绝了我们最初的组合链;修复后的链通过两个已证明的桥接跨越度量,任何形式系统无法验证的内容则改为测量,自动将组合层降至经验层:所有主张要么已验证、部分验证,要么为经验,组合继承最弱的层,且该层由机器决定。在1240万次条目读取上重放,在八路并发、每请求预算及故障关闭的身份归因下运行,该账本量化了当前见证的诚实权衡,并在零违规情况下维持其风险预算。一个融合的始终在线探针在服务噪声底限内的CUDA图下观测声明的单层子集。应用于带有打包压缩KV原型的已部署DeepSeek-V4栈,相同机制通过机器裁定的判别活动将静默故障定位到精确的结构边界——在无驱逐、身份隔离的区域中精确,在驱逐或槽复用区域中观测到的每一次故障亦然——此过程中,该判别活动的演算拒绝了我们自身的两个混淆推断。所有工件、保护措施及Lean开发内容均已在该httpsURL发布,本文中每个数字均可通过一个命令从已发布工件中重新生成。

英文摘要

Modern models no longer keep a plain KV cache: latent caches, learned sparse selectors and recurrent states each carry the model's memory in a different form, and each fails differently under compression. We give a runtime observability contract that covers all four memory classes with three operators, instantiate it on six model configurations across five architecture families, and compose the per-stage bounds into an executable request-level risk ledger. Contracts carry their error metric as a type -- composition is only defined when metrics match, and this check rejected our own first composed chain; the repaired chain crosses metrics through two proved bridges, and whatever no formal system can certify is measured instead, dropping the composed tier to empirical automatically: every claim is certified, partially certified, or empirical, composition inherits the weakest tier, and the tier is decided by the machine. Replayed over $12.4$M entry reads and run under eight-way concurrency with per-request budgets and fail-closed identity attribution, the ledger quantifies the honest trade-off on today's witness and holds its risk budget with zero violations. A fused always-on probe observes a declared one-layer subset under CUDA graphs inside the serving noise floor. Applied to a served DeepSeek-V4 stack with a packed compressed-KV prototype, the same machinery localizes a silent corruption to a precise structural boundary -- exact in the eviction-free, identity-isolated regime, with every observed failure in an eviction or slot-reuse regime -- through a machine-adjudicated discrimination campaign whose calculus rejected two of our own confounded inferences along the way. All artifacts, guards, and the Lean development are released at https://github.com/metask-ai/witprobe-attention-memory; every number in this paper regenerates from the shipped artifacts by one command.

补充信息

↑