arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Hoss:基于异构GPU-CPU-TEE架构的快速不经意语义搜索

Hoss: Fast Oblivious Semantic Search with Heterogeneous GPU-CPU-TEE Architecture

Jianzhang Du, Weijie Huang, Chenghong Wang, Nicolas Tsagareli, Yukui Luo, XiaoFeng Wang, Zhongshu Gu

arXiv 2609.04522首次发表:更新:

发表机构

Indiana University; Binghamton University; Nanyang Technological University; IBM Research(印第安纳大学; 宾汉姆顿大学; 南洋理工大学; IBM研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对现有不经意语义搜索系统Compass开销大的问题,提出基于异构CPU-GPU TEE架构的Hoss系统,利用GPU TEE的大容量Pmem优化,实现最高67倍加速且保持高召回率。

AI 中文摘要

语义搜索已广泛部署于现代AI系统中,但保护数据内容与访问模式仍具挑战性。当前最先进的系统Compass通过在HNSW图上构建优化的ORAM实现不经意语义搜索,不过即便采用激进优化,其仍存在较大开销。缩小这一性能差距从根本上存在困难:Compass已消除大部分加密开销,使得ORAM访问成为主导成本,而ORAM访问受限于已知的Omega(log N)带宽下界。本文核心洞见在于,传统ORAM开销源于有限私有内存的假设,而现代GPU TEE提供可屏蔽内部访问模式的大容量私有内存(Pmem)(Hunt等人,NSDI '23),这一转变开辟了新的设计空间。因此,我们提出Hoss,这是首个基于异构CPU-GPU TEE架构的不经意语义搜索系统,支持快速、可扩展且拥有低所有权成本的搜索。在Hoss中,GPU TEE的大容量Pmem承载热路径HNSW遍历,若图的下层超出GPU容量则卸载至CPU TEE,系统仅在访问这些下层时调用不经意原语。大容量Pmem的可用性还带来新的优化机会,例如Hoss具备超出传统性能约束的主机访问ORAM机制,并整合了先前设计无法实现的多种依赖数据的优化。我们实现了Hoss原型并与Compass进行基准测试,结果显示Hoss在保持高召回率的同时实现了最高达67倍的加速,且规模越大增益越显著。

英文摘要

Semantic search is widely deployed in modern AI systems, but protecting both data contents and access patterns remains challenging. The current state-of-the-art system, Compass, achieves oblivious semantic search by building an optimized ORAM over HNSW graphs. However, even with aggressive optimizations, it still incurs large overheads. Closing this performance gap is fundamentally difficult: Compass has already removed most cryptographic overheads, leaving ORAM accesses as the dominant cost, which are constrained by well-known Omega(log N) bandwidth lower bounds. Our key insight is that traditional ORAM overhead stems from the assumption of limited private memory, whereas modern GPU TEEs provide large private memory (Pmem) that blinds internal access patterns (Hunt et al., NSDI '23). This shift opens a new design space. We therefore propose Hoss, a first-of-its-kind oblivious semantic search system with a heterogeneous CPU-GPU TEE architecture that supports fast, scalable search with low cost of ownership. In Hoss, the GPU TEE's large Pmem hosts the hot-path HNSW traversal, while the lower layers of the graph, if they exceed GPU capacity, are offloaded to CPU TEEs. The system invokes oblivious primitives only when accessing these lower layers. The availability of large Pmem also enables new optimization opportunities. For example, Hoss features a host-access ORAM mechanism that goes beyond traditional performance constraints and incorporates several data-dependent optimizations that are not possible in prior designs. We implement a prototype of Hoss and benchmark it against Compass. Our results show that Hoss achieves up to 67x speedup while maintaining high recall, with larger gains at scale.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑