arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

KumbhDoot:面向大规模集会公共服务助手的可扩展、大语言模型(LLM)限定架构

KumbhDoot: A Scale-Ready, LLM-Bounded Architecture for Mass-Gathering Public-Service Assistants

Saurabh Sakalkar, Abhishek Singh, Ramesh Raskar

arXiv 2608.07520首次发表:更新:

发表机构

Cisco; MIT(思科; 麻省理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对大壶节这类高风险低连接性的大规模集会,提出KumbhDoot架构,以相似度优先、LLM限定为核心,解决默认LLM助手成本高、易幻觉等问题,适配公共服务场景。

AI 中文摘要

Kumbh Mela(大壶节)这类大规模宗教集会在数周内将数千万人集中到单一区域,产生密集、重复、多语言且关乎安全的信息需求。默认方案是将所有查询路由到大语言模型(LLM)的对话助手,该方案在此场景下适配性差:大规模部署成本高、紧急路径响应慢、易出现可能造成人身伤害的事实幻觉,且无网络时无法使用。本文介绍KumbhDoot,这是为Nashik Simhastha Kumbh Mela打造的智能朝圣者助手,遵循不同原则:其基础设计原则优先考虑语义相似度,而非从LLM出发;仅当基于相似度的检索无法生成正确答案时,才调用生成式模型。该系统采用“语义缓存”:以嵌入索引存储作为单一检索原语,处理意图路由、答案缓存、离线查询及多智能体检索;自定义三层智能体架构直接基于该存储运行,确保决策路径可检查,避免使用通用多智能体框架(会触发隐式的每步LLM调用)。本文呈现该架构、其单查询经济性的分析成本模型,以及关于相似度何时足够、生成式推理仍有必要的客观说明;本文认为,对于受限、高风险、低连接性的公共服务领域,相似度优先且LLM限定的设计不仅成本更低,在架构上也比默认采用LLM的方案更合适。

英文摘要

Mass religious gatherings such as the Kumbh Mela concentrate tens of millions of people into a single region over a few weeks, producing intense, repetitive, multilingual, and safety-critical demand for information. The default response, a conversational assistant that routes every query to a large language model (LLM), is poorly matched to this setting: it is costly at scale, slow on emergency paths, prone to hallucination on facts that can cause physical harm, and unusable when connectivity fails. We describe KumbhDoot, an agentic pilgrim assistant for the Nashik Simhastha Kumbh Mela built on a different principle. It operates on a foundational design principle that prioritizes semantic similarity over starting with an LLM. Generative models are invoked only in instances where similarity-based retrieval is insufficient to produce a correct answer. The system utilizes a "semantic cache": an embedding-indexed store as a single retrieval primitive, which handles intent routing, answer caching, offline lookups, and multi-agent retrieval. A custom three-tier agent architecture operates directly on this store, ensuring decision paths remain inspectable and avoiding the use of generic multi-agent frameworks that would trigger implicit per-step LLM calls. We present the architecture, an analytical cost model for its per-query economics, and an honest account of where similarity is sufficient and where generative reasoning remains necessary. We argue that for bounded, high-stakes, low-connectivity public-service domains, a similarity-first and LLM-bounded design is not merely cheaper but architecturally more appropriate than an LLM-default one.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑