PULSAR:面向企业视觉文档RAG的池化统一后期交互搜索与检索
PULSAR: Pooled Unified Late-Interaction Search and Retrieval for Enterprise Visual Document RAG
AI总结:
针对企业视觉文档RAG的检索难题,提出PULSAR系统,采用池化两阶段后期交互索引,在延迟、QPS、成本等指标上显著优于OCR基线,已在生产环境大规模应用。
AI中文摘要:
机构投资者搜索视觉密集型的演示文稿、董事会材料和尽职调查材料,这些材料在交易收尾阶段每小时都会变化。光学字符识别(OCR)后接图表文字化处理,在这种规模下刷新成本高昂,且会丢失图表细节。我们提出PULSAR,这是一款部署于穆巴达拉投资公司(Mubadala Investment Company)的生产级以视觉为核心的检索系统。PULSAR采用冻结的ColPali风格主干网络为页面图像建立索引,并使用池化的两阶段后期交互索引:紧凑的页面摘要支持初始检索,随后在更精细的池化表示上进行精确的最大相似度(MaxSim)重排序。在ViDoRe V3数据集上,该设计相较于未池化配置,将向量搜索的中位数延迟降低了15.1倍,同时NDCG@10和Recall@10的绝对损失小于0.01;生产环境下的向量搜索中位数延迟为156毫秒。在并发负载下,池化索引的每秒查询数(QPS)比未池化索引高出约88倍。据估算,事件驱动的摄取路径每页成本约为其取代的OCR+文字化基线的1/20。自2026年3月以来,PULSAR已为超过3000笔交易提供了7.8万份文档和约240万页内容,在生产环境的Top K设置下,其答案事实召回率较OCR+文字化基线提升了一倍以上。
英文摘要:
Institutional investors search visually dense pitch decks, board packs, and diligence materials that change hourly near deal closing. OCR followed by figure verbalisation is costly to refresh at this scale and can lose chart detail. We present PULSAR, a production vision-first retrieval system deployed at Mubadala Investment Company. PULSAR indexes page images with a frozen ColPali-style backbone and uses a pooled two-stage late-interaction index: compact page summaries support initial retrieval, followed by exact MaxSim rescoring over a finer pooled representation. On ViDoRe V3, this design reduces median vector-search latency by 15.1 times against an unpooled configuration with less than 0.01 absolute NDCG@10 and Recall@10 loss; production median vector-search latency is 156 ms. Under concurrent load, the pooled index sustains approximately 88 times higher QPS than an unpooled index. The event-driven ingestion path is estimated to be approximately 20 times cheaper per page than the OCR+verbalisation baseline it replaced. Since March 2026, PULSAR has served 78 thousand documents and approximately 2.4 million pages across more than 3,000 deals. At the production top K, it more than doubles answer-fact recall over the OCR+verbalisation baseline.