AI 中文总结
该研究推出用于评估大型语言模型伊斯兰学术性能的多任务基准 ISTB,含3465个问答条目,支持多维度可复现评估,填补了该领域高质量标注资源的空白。
AI 中文摘要
大型语言模型(LLMs)越来越多地被用于问答、教育和研究,包括在答案依赖专业来源传统的宗教和文化领域。然而,在伊斯兰研究中,被称为 turath 的权威学术传统中保存的关键概念、方法和辩论,缺乏高质量的标注资源。我们推出 IslamicTurathBench(ISTB),这是一个用于评估大型语言模型在古典伊斯兰学术领域性能的多任务、多学科数据集。ISTB 由领域专家开发和审核,包含来自 35 部公认来源著作的 3465 个问答条目,这些著作涵盖了伊斯兰研究七个关键领域、超过 12 个世纪的学术成果。为了全面分析模型能力,ISTB 沿两个维度构建:学术需求(初级、中级和高级)和任务形式(多项选择题、基于段落的理解题以及开放式知识题)。ISTB 包含学术人类参考小组的汇总分数,以及来自十个系统的零样本基线。该数据集支持在具有历史分层的学术领域中,针对来源著作、学科、学术需求水平和问题形式的语言模型行为的可复现评估。
英文摘要
Large language models (LLMs) are increasingly used for question answering, education, and research, including in religious and cultural domains where answers depend on specialised source traditions. Yet in Islamic Studies, key concepts, methods, and debates preserved in the authoritative scholarly tradition, known as turath, lack high-quality annotated resources. We introduce IslamicTurathBench (ISTB), a multi-task, multi-discipline dataset for evaluating LLMs on classical Islamic scholarship. Developed and reviewed by domain experts, ISTB contains 3,465 question-answer items drawn from 35 recognised source works spanning more than 12 centuries of scholarship across seven key fields of Islamic Studies. To enable comprehensive profiling of model capabilities, ISTB is structured along two axes: scholarly demand (Beginner, Intermediate, and Advanced) and task format (multiple-choice questions, passage-based comprehension, and open-ended knowledge questions). ISTB includes aggregated scores from a scholarly human reference panel and zero-shot baselines from ten systems. The dataset supports reproducible evaluation of language-model behaviour across source works, disciplines, scholarly-demand levels, and question formats in a historically layered scholarly domain.
CommentsIncludes supplementary materials. Submitted to the Journal of Scientific Data. Data and code are publicly available