arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.10144cs.AIcs.LG

内核管理的共享内存:实现系统级个性化

Kernel-Managed Shared Memory for System-Wide Personalization

Ryan Lum, Yongfeng Zhang

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出内核管理的共享内存,由系统内核统一管理多智能体的记忆检索与隐私,在AIOS上测试显示个性化评分显著提升且延迟降低15-61%,成本更低。

中文摘要 AI 辅助

当人工智能系统能够适应使用它们的人群时,它们会变得更加有用,但在多智能体系统中,一个智能体学到的有用上下文通常对其他智能体不可用。我们提出了内核管理的共享内存,这是一种系统级抽象,其中专门的智能体写入结构化的、带标签的记忆,而智能体系统内核(而非单个智能体)负责管理检索、隐私执行和提示注入。我们在AIOS上实现并评估了这种设计,并将其与三种替代方案在三个助手模型(GPT-4o、Llama-3.1:8B、Qwen-2.5:7B)和总共1800次试验中进行了比较。与使用相同底层存储的未管理外部记忆后端(Mem0)相比,内核管理的检索和注入在5分制上将个性化评分提高了2.4-4.0分(例如,GPT-4o上的档案使用率从1.05提高到4.69),每次比较在p < 10^-18时均具有显著性。与标准的检索增强注入相比,增益同样显著且在所有三个模型中保持一致。与完整的、未过滤的上下文拼接(这是可用上下文而非响应质量的软上限)相比,内核管理的注入在三个模型中的两个上统计上匹配了性能,在第三个上显示出小的、特定于模型的缺陷,同时使用了显著更短的提示:所有三个模型的端到端延迟降低了15-61%,相应的每次调用令牌使用量和推理成本也有所降低。这些结果表明,将记忆管理集中到智能体系统内核中,而不是将检索和隐私执行留给单个智能体,以一小部分成本实现了无约束上下文的大部分个性化收益。

英文摘要

AI systems become more useful when they can adapt to the people using them, but in multi-agent systems, useful context learned by one agent often remains unavailable to others. We present kernel-managed shared memory, a system-level abstraction in which specialized agents write structured, tagged memories while the agent-system kernel, not individual agents, governs retrieval, privacy enforcement, and prompt injection. We implement and evaluate this design on AIOS and compare it against three alternatives across three assistant models (GPT-4o, Llama-3.1:8B, Qwen-2.5:7B) and 1,800 total trials. Against an unmanaged external memory backend (Mem0) using identical underlying storage, kernel-managed retrieval and injection improve personalization scores by 2.4-4.0 points on a 5-point scale (e.g., 1.05 to 4.69 profile usage on GPT-4o), with every comparison significant at p < 10^-18. Against standard retrieval-augmented injection, gains are similarly large and consistent across all three models. Against full, unfiltered context concatenation, a soft ceiling on available context rather than on response quality, kernel-managed injection statistically matches performance on two of three models and shows a small, model-specific deficit on the third, while using substantially shorter prompts: end-to-end latency is 15-61% lower across all three models, with corresponding reductions in per-call token usage and inference cost. These results indicate that centralizing memory management in the agent-system kernel, rather than leaving retrieval and privacy enforcement to individual agents, delivers most of the personalization benefit of unconstrained context at a fraction of its cost.

发表机构

  • Rutgers University(罗格斯大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑