发表机构
Greentech Apps Foundation; Queen Mary University of London(绿色科技应用基金会; 伦敦玛丽女王大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文通过批判性叙事综述,梳理大型语言模型等技术对圣训计算科学的重塑进展,指出其在数据、任务等方面的成果,同时分析基准、语料库等局限,提出需将其作为证据基础设施问题研究的议程。
AI 中文摘要
本文探讨了大型语言模型(LLMs)、检索增强流程及Transformer模型如何重塑圣训计算科学。现有综述记录了该领域文献的增长,但未批判性梳理哪些进展在方法学上稳健、哪些仍局限于基准测试、哪些未解决问题仍制约学术应用。本文通过批判性叙事综述填补这一空白,该综述结合对现有综述的批判、对代表性原创研究的论文级评估,以及伊斯兰学者和领域专家对真实性、权威性及负责任使用的观点综合。研究发现进展不均衡:数据资源已扩展,分段任务已成熟,叙述者与来源验证问题得到更好形式化,LLM辅助工作流现已支持语料库规模的丰富化、多语言访问及基于基准的评估。与此同时,进展仍受限于语料库范围狭窄、基准可比性弱、合成数据到真实数据的迁移差距、叙述者身份解析、预处理脆弱性、可复现性有限及基于专家的验证不足。研究表明,重要缺口存在于主流基准之外:非正典及晦涩语料库、评注与解释性文献、与《古兰经》和圣行(seerah)的跨来源链接、以及面向伊斯兰教法(fiqh)的证据支持。本文主张,圣训计算的评估不应仅视为孤立的模型性能问题,而应视为需要知识整合、溯源及专家监督的证据基础设施问题。基于此,本文定义了一项研究议程,以增强该领域的方法学稳健性并提升其对伊斯兰学术的实用性。
英文摘要
We examine how hadith computational science is being reshaped by transformer models, retrieval-grounded pipelines, and large language models (LLMs). Recent reviews document growth in the literature, but they do not yet provide a critical account of which advances are methodologically robust, which remain benchmark-bound, and which unresolved problems still limit scholarly use. We address this gap through a critical narrative review that combines critique of existing reviews, paper-level appraisal of representative original studies, and synthesis of Islamic scholar and domain-expert perspectives on authenticity, authority, and responsible use. We find uneven progress. Data resources have expanded, segmentation tasks have matured, narrator and source-verification problems are better formalized, and LLM-assisted workflows now support corpus-scale enrichment, multilingual access, and grounded evaluation. At the same time, progress remains constrained by narrow corpora, weak benchmark comparability, synthetic-to-real transfer gaps, narrator identity resolution, preprocessing fragility, limited reproducibility, and sparse expert-grounded validation. We show that important gaps lie beyond dominant benchmarks: non-canonical and obscure corpora, commentary and explanatory literature, cross-source links with Qur'an and seerah, and fiqh-facing evidence support. We argue that hadith computation should be assessed less as isolated model performance than as an evidence infrastructure problem requiring knowledge integration, provenance, and expert supervision. On this basis, we define a research agenda for making the field methodologically stronger and more useful to Islamic scholarship.
CommentsSubmitted to Artificial Intelligence Review