AI 中文总结
提出Zengram-Lite,一个在浏览器中运行的智能体记忆框架,提供知识、会话跟踪和上下文组装三层,通过六阶段组装实现令牌预算提示,并继承事务一致性。
AI 中文摘要
人工智能智能体越来越多地在浏览器中运行,它们需要存储所学内容的地方。然而,目前的客户端状态技术是向量索引——对嵌入进行最近邻搜索——而智能体记忆的其余部分则留给应用程序代码:对话历史是localStorage中的一个数组,上下文管理是手工处理的截断,并且没有关于事实置信度、来源或产生它的会话的共享概念。我们提出zengram-lite,一个编译为单个约2.95 MB gzipped WebAssembly工件、完全在浏览器标签页中运行的智能体记忆框架。它在单个事务存储之上提供三个层级。知识层提供混合向量和全文检索,具有事实生命周期——置信度随事实被确认或反驳而上升或下降,重要性随时间衰减,取代将重述的事实视为更新,以及作用域命名空间。会话跟踪层将对话本身建模为一等数据:会话、轮次、类型化内容部分以及带有状态和时机的工具调用,全部作为可查询表。上下文组装层通过六阶段组装将该结构转换为令牌预算的提示,在调用方的预算下打包系统指令、知识和近期历史,并返回一个稳定的指纹,指示何时可以重用提示前缀。该捆绑包是zeta-lite SQL引擎的超集——同一个.wasm重新导出完整的Postgres兼容表面、MVCC快照隔离和写时复制数据库分支——因此记忆在所有三个层级中继承事务一致性。由于wasm表面是同步的,而框架的规范操作依赖于异步LLM和嵌入器,zengram-lite暴露了一个自带结果接缝,该接缝在应用程序在JavaScript中计算的结果上运行框架的真实代码路径。
英文摘要
AI agents increasingly run in the browser, and they need somewhere to keep what they learn. The client-side state of the art, however, is a vector index - nearest-neighbor search over embeddings - with the rest of an agent's memory left to application code: the conversation history is an array in localStorage, context management is hand-rolled truncation, and there is no shared notion of a fact's confidence, its provenance, or the session that produced it. We present zengram-lite, an agentic-memory framework compiled to a single ~2.95 MB gzipped WebAssembly artifact that runs entirely in a browser tab. It provides three tiers over one transactional store. The knowledge tier offers hybrid vector-and-full-text recall with a fact lifecycle - confidence that rises and falls as facts are confirmed or contradicted, importance that decays over time, supersession that treats a restated fact as an update, and scope namespacing. The session-tracking tier models the conversation itself as first-class data: sessions, turns, typed content parts, and tool calls with state and timing, all as queryable tables. The context-assembly tier turns that structure into a token-budgeted prompt through a six-phase assembly that packs system instructions, knowledge, and recent history under a caller's budget, and returns a stable fingerprint that signals when a prompt prefix can be reused. The bundle is a superset of the zeta-lite SQL engine - the same .wasm re-exports the full Postgres-compatible surface, MVCC snapshot isolation, and copy-on-write database branching - so memory inherits transactional consistency across all three tiers. Because the wasm surface is synchronous while the framework's canonical operations depend on an asynchronous LLM and embedder, zengram-lite exposes a bring-your-own-result seam that runs the framework's real code paths over results the application computes in JavaScript.