arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

自进化搜索索引

Self-Evolving Search Index

Sangam Lee, Wonjae Lee, Sunghwan Kim, Deogyong Kim, Jaehoon Kim, Daye Nam, SeongKu Kang, Dongha Lee

arXiv 2609.19656首次发表:更新:

发表机构

Yonsei University; Samsung Research; University of California, Irvine; Korea University(延世大学; 三星研究院; 加州大学欧文分校; 高丽大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对索引优化依赖人工的问题,提出SELF-INDEX框架,通过优化器自主诊断、修订并验证索引键,结合查询模拟器主动探索需求,实现索引自我进化,持续提升检索性能并惠及下游应用。

AI 中文摘要

随着LLM智能体处理涉及多样化信息需求的复杂任务,信息检索变得越来越重要。由于检索依赖于通过索引键表示每个文档的索引,检索质量在很大程度上取决于这些键如何有效地揭示每个文档中包含的知识。然而,有效的索引表示因检索环境而异,这使得任何固定的优化策略都难以保持一致的表现。但将索引进化到其检索环境仍然主要依赖人工驱动,需要人类诊断检索失败、优化优化策略并相应地重新处理索引。我们提出SELF-INDEX,一个使索引无需人工干预即可自我进化的框架。其优化器自主诊断检索不足,选择性地修订负责的索引键,并在更新索引前验证每次修订。除了对观察到的检索需求做出反应外,SELF-INDEX还通过查询模拟器主动探索额外需求,使索引能够超越优化时已有的查询进行进化。在多样化的语料库和检索器中,SELF-INDEX持续提升检索性能,并优于现有的索引优化方法。我们进一步表明,这些益处扩展到下游应用,提高了搜索智能体的有效性和效率,并帮助智能体记忆系统检索有用的过往交互。

英文摘要

Information retrieval is increasingly important as LLM agents tackle complex tasks involving diverse information needs. Because retrieval relies on an index that represents each document through index keys, retrieval quality depends heavily on how effectively these keys expose the knowledge contained in each document. However, effective index representations vary across retrieval environments, making it difficult for any fixed optimization strategy to perform consistently. Yet evolving an index to its retrieval environment remains largely human-driven, requiring humans to diagnose retrieval failures, refine the optimization strategy, and reprocess the index accordingly. We propose SELF-INDEX, a framework that enables an index to self-evolve without human intervention. Its Optimizer autonomously diagnoses retrieval shortfalls, selectively revises the responsible index keys, and validates each revision before updating the index. Beyond reacting to observed retrieval demands, SELF-INDEX proactively explores additional demands through a Query Simulator, allowing the index to evolve beyond the queries already available for optimization. Across diverse corpora and retrievers, SELF-INDEX consistently improves retrieval performance while outperforming existing index optimization methods. We further show that these benefits extend to downstream applications, improving the effectiveness and efficiency of search agents and helping agent memory systems retrieve useful past interactions.

CommentsWork in progress

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑