发表机构
Xiaohongshu(小红书)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对企业数据分析智能体检索数据资产的失效问题,提出两层解决方案,经小红书广告数据仓库验证,可大幅提升检索准确率与知识覆盖率,降低延迟。
AI 中文摘要
企业数据分析智能体面临两类结构性失效:通用检索增强生成(RAG)会检索到错误资产(Hit@10=19.1%),且无法提供使用知识以避免指标误读,其根源在于语义鸿沟、实体歧义、模式漂移和资产-使用缺口四类根本原因(C1--C4)。我们在小红书商业广告数据仓库(含5300+个Hive表,覆盖14个领域)部署了两层解决方案:三层双用途知识库(含179份文档,采用八部分标注模板)同时服务于检索和生成,闭环刷新管道维持日级时效性(仅需一次是/否审批,30秒热重载);图引导检索器(GGR)以2859节点的知识图谱作为候选门,结合意图路由实现71.6倍的令牌减少;场景感知排序器(SAR)应用19类实体识别和显式场景标注,仅负知识一项就带来Hit@10提升25个百分点。在两个含100个问题的基准测试中,Hit@10从19.1%升至96.6%(提升77.5个百分点),知识覆盖率从56%升至77%,端到端延迟为4.84--5.33秒。
英文摘要
Enterprise data analytics agents face two structural failures: generic RAG retrieves the wrong asset (Hit@10=19.1%) and delivers no usage knowledge to prevent metric misinterpretation---stemming from four root causes (C1--C4) ranging from semantic gap and entity ambiguity to schema drift and asset-usage gap. We present a two-layer solution deployed in the commercial advertising data warehouse at Xiaohongshu (5,300+ Hive tables, 14 domains). A three-tier dual-purpose knowledge base (179 documents, eight-section annotation template) serves both retrieval and generation, with a closed-loop refresh pipeline maintaining day-level freshness (one yes/no approval, 30s hot-reload). The Graph-Guided Retriever (GGR) uses a 2,859-node knowledge graph as a candidate gate with intent routing to deliver 71.6x token reduction. The Scene-Aware Ranker (SAR) applies 19-class entity recognition and explicit scenario annotations; negative knowledge alone contributes 25 percentage points of Hit@10 gain. On two 100-question benchmarks, Hit@10 rises from 19.1% to 96.6% (+77.5pp) and knowledge coverage from 56% to 77%, at 4.84--5.33s end-to-end latency.
Comments6 pages, 2 figures, 2 tables. Submitted to DAI 2026 Industry Track