arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2606.29652cs.IR

如我们可能搜索

As We May Search

Saber Zerhoudi, Adam Roegiest, Jelena Mitrovic, Michael Granitzer

更新

AI总结:

针对个人文档等敏感信息搜索中的隐私问题,提出本地优先信息检索(local-first IR)设计理念,通过将索引、模型和推理置于用户设备上,在消费级硬件上实现与云端相当的性能,并指出真正的权衡在于搜索范围而非质量。

AI中文摘要:

个人文档、法律文件和医疗记录中的敏感信息是最有价值的搜索对象之一,但当前的检索增强生成系统仍然需要将内容发送到远程服务器。我们提出本地优先信息检索(local-first IR),这是一种设计理念,其中索引、模型和推理位于用户设备上,将远程服务视为可选项。本文做出四项贡献:(1) 一个沿三个维度组织检索架构的框架:隐私与控制、能力、可访问性;(2) 在消费级硬件上跨五个基准的实验,从1K到1M文档,使用稠密检索、BM25和混合融合。稠密检索在10万文档内保持超过91%的nDCG@10,近似HNSW索引将其扩展到100万文档,质量损失仅为2%;一个7B本地语言模型在答案质量上达到与云基线相差4个点以内;(3) 基于实验证据的支持和反对本地优先IR的竞争视角;(4) 一个识别开放问题的研究议程。真正的权衡是范围而非质量:重要的是你能搜索什么,而不是你能搜索得多好。

英文摘要:

The sensitive information in personal documents, legal files, and medical records is among the most valuable things to search, yet current retrieval-augmented generation systems still require sending content to remote servers. We propose local-first IR, a design philosophy where indexes, models, and inference reside on user devices, treating remote services as optional. This paper makes four contributions: (1) a framework organizing retrieval architectures along three dimensions: privacy and control, capability, and accessibility, (2) experiments on consumer hardware across five benchmarks, scaling from 1K to 1M documents with dense retrieval, BM25, and hybrid fusion. Dense retrieval keeps over 91% nDCG@10 up to 100K documents, with approximate HNSW indexes extending this to 1M with only 2% quality loss; a 7B local language model reaches within 4 points of a cloud baseline on answer quality, (3) competing perspectives for and against local-first IR, informed by experimental evidence, and (4) a research agenda identifying open problems. The real tradeoff is scope rather than quality: what matters is what you can search, not how well you can search it.

补充信息

↑