arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.23371cs.CLcs.LG

机器可解释信息:将文档编译为可搜索且可读的协议状态

Machine-Interpretable Information: Compiling Documents into Searchable and Readable Protocol States

Yifan Wang, Dejing Dou

首次发表
浏览论文内容

中文总结 AI 辅助

提出MII协议,通过双时间尺度Writer将文档编译为固定带宽状态,实现跨模型可转移的检索、推理与重建,Residual-MII在HotpotQA上以7%计算成本超越全上下文精确匹配。

中文摘要 AI 辅助

长上下文语言模型通过原始自然语言与外部知识交互。在检索增强系统中,这造成了持续的索引-载荷分裂:密集向量实现了可搜索的路由,但模型必须重新读取冗长的文本载荷以进行推理,代价是O(N^2)的注意力计算。现有压缩方法进一步产生了与特定架构绑定的私有状态。我们引入了机器可解释信息(MII),这是首个智能体到智能体(A2A)的文档到状态协议。一个双时间尺度的状态空间Writer将文档编译为规范的、固定带宽的状态(56个token),而一个轻量级Translator将其映射到任何冻结的Reader的嵌入空间,将查询时成本降低到O(K)。生成的.mii工件在单一可转移介质中统一了检索(可搜索的几何结构)、推理(全局记忆)和重建(基于事实的细节)。我们展示了跨异构LLM(如Llama、Qwen、Mistral)的强跨模型互操作性——尽管Writer使用旧版GPT-2词汇表,迫使进行真正的语义翻译而非token级记忆。机制探针揭示了模块化的潜在结构:实体表示可以被因果追踪,并在不相关的文档状态之间零样本移植,同时保持可解码性。为了解决固定带宽下的词汇重建问题,我们提出了Residual-MII,一种结合编译全局记忆与稀疏局部证据的缓存层次结构。在HotpotQA(7,405个查询)上,Residual-MII在约7%的注意力FLOPs下超过了全上下文精确匹配,表明向编译的、可转移的神经文档格式的范式转变。

英文摘要

Long-context language models interface with external knowledge through raw natural language. In retrieval-augmented systems, this creates a persistent index-payload schism: dense vectors enable searchable routing, but models must re-ingest lengthy text payloads for reasoning at O(N^2) attention cost. Existing compression methods further produce private states tied to specific architectures. We introduce Machine-Interpretable Information (MII), the first agent-to-agent (A2A) document-to-state protocol. A dual-timescale state-space Writer compiles documents into a canonical, fixed-bandwidth state (56 tokens), and a lightweight Translator maps it into any frozen Reader's embedding space, reducing query-time cost to O(K). The resulting .mii artifact unifies Retrieval (searchable geometry), Reasoning (global memory), and Reconstruction (grounded details) in a single transferable medium. We demonstrate strong cross-model interoperability across heterogeneous LLMs (e.g., Llama, Qwen, Mistral) -- despite the Writer using a legacy GPT-2 vocabulary, forcing genuine semantic translation rather than token-level memorization. Mechanistic probes reveal modular latent structure: entity representations can be causally traced and zero-shot transplanted between unrelated document states while remaining decodable. To address lexical reconstruction under fixed bandwidth, we propose Residual-MII, a cache hierarchy combining compiled global memory with sparse local evidence. On HotpotQA (7,405 queries), Residual-MII exceeds full-context Exact Match at approximately 7% of the attention FLOPs, suggesting a paradigm shift toward compiled, transferable neural document formats.

发表机构

  • NexusLumenLabs
  • Fudan University(复旦大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑