arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.20390cs.CLcs.AIcs.CY

Ansari:基于检索的伊斯兰AI助手——架构、部署及14万次对话的经验教训

Ansari: A Retrieval-Grounded Islamic AI Assistant -- Architecture, Deployment, and Lessons from 140,000 Conversations

M Waleed Kadous, Amr Elsayed, Abdullah Al Nahas, Ashraf Haress

首次发表
浏览论文内容

中文总结 AI 辅助

本研究推出基于检索的伊斯兰AI助手Ansari,其经多平台部署后处理14万次跨语言对话,在IslamicMMLU等基准中表现优异,为信仰敏感型LLM部署提供了经验。

中文摘要 AI 辅助

通用大型语言模型(LLM)越来越多地用于回答宗教问题,但针对伊斯兰内容,它们存在两大严重风险:事实编造(编造《古兰经》经文或圣训)和微妙的价值对齐偏差。我们推出Ansari,这是一个已部署的、基于检索的伊斯兰AI助手,自2023年6月以来已处理超过14万次跨25种以上语言的对话。Ansari围绕智能体检索循环构建:使用工具的语言模型对经认证的伊斯兰语料库进行搜索——包括《古兰经》、圣训集、多卷本教法(fiqh)百科全书和经注(tafsir)来源——且仅基于检索到的内容作答,并附上引文以供验证。我们描述了系统的架构(智能体循环、检索工具、语料库以及编码编辑和神学政策的系统提示)、其多平台部署(网页、移动应用、WhatsApp,以及作为模型上下文协议服务器和智能体技能),以及14万次真实对话揭示的穆斯林实际使用此类工具的方式。我们报告了多项补充评估的结果:在认可的机构考试上的零样本性能、斋月期间的人工评分验证,以及两个独立的外部运行基准,在这些基准中,Ansari目前在公开的IslamicMMLU排行榜上领先于前沿模型,在伊斯兰法律推理(IslamicLegalBench)上具有竞争力,同时强烈抵制错误前提——并得出了适用于伊斯兰教之外任何信仰或价值敏感型LLM部署的经验教训:基础是必要但不充分的,系统提示既是技术产物也是神学产物,模型构建过程中缺乏社区参与仍是一个难以解决的缺口。

英文摘要

General-purpose large language models (LLMs) are increasingly used to answer religious questions, but for Islamic content they carry two serious risks: factual fabrication (inventing Qur'anic verses or hadith) and subtle value misalignment. We present Ansari, a deployed, retrieval-grounded Islamic AI assistant that has handled more than 140,000 conversations across 25+ languages since June 2023. Ansari is built around an agentic retrieval loop: a tool-using language model issues searches against authenticated Islamic corpora -- the Qur'an, hadith collections, a multi-volume jurisprudence (fiqh) encyclopedia, and exegetical (tafsir) sources -- and answers only on the basis of what it retrieves, with citations attached for verification. We describe the system's architecture (the agent loop, the retrieval tools, the corpora, and the system prompt that encodes editorial and theological policy), its multi-platform deployment (web, mobile, WhatsApp, and as a Model Context Protocol server and an Agent Skill), and what 140,000 real conversations reveal about how Muslims actually use such a tool. We report results on several complementary evaluations -- zero-shot performance on accredited institutional exams, a human-rated validation during Ramadan, and two independent, externally run benchmarks on which Ansari currently tops the public IslamicMMLU leaderboard ahead of frontier models and is competitive on Islamic legal reasoning (IslamicLegalBench) while strongly resisting false premises -- and draw out lessons that generalize beyond Islam to any faith- or values-sensitive deployment of LLMs: grounding is necessary but not sufficient, the system prompt is a theological as much as a technical artifact, and the absence of community in how models are formed remains a hard gap.

补充信息

↑