arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.20246cs.IR

什么样的伊斯兰教法检索器是优秀的?阿拉伯伊斯兰教法的答案检索

What Makes a Good Fiqh Retriever? Answer Retrieval for Arabic Islamic Jurisprudence

Somaya Eltanbouly, Heba Sbahi, Samer Rashwani, Abdessalam Bouchekif, Mutaz al-Khatib, Shahd Gaben, Mohammed Ghaly

首次发表
浏览论文内容

中文总结 AI 辅助

本研究针对阿拉伯伊斯兰教法构建检索测试集,评估多种检索策略,发现教法学派感知过滤可显著提升教派特定问题的检索性能,核心挑战是区分答案承载与主题相似的非答案段落。

中文摘要 AI 辅助

检索增强生成技术被用于伊斯兰问答任务,但大多数系统采用端到端评估,难以区分检索失败与生成失败。本研究针对阿拉伯伊斯兰教法(fiqh)开展答案承载段落的检索研究,定义仅当段落陈述问题所需的裁决(ruling)时才为相关。我们构建了阿拉伯伊斯兰教法的检索测试集,用于评估密集型、词汇型、混合型、微调型以及教法学派(madhhab)感知的检索策略。最优检索器的MRR@5达0.524,微调后性能提升至0.553;混合型检索对强模型增益有限,而教法学派感知过滤在教派特定问题上使MRR@5提升超一倍。进一步的错误分析显示,核心挑战是区分答案承载段落与主题相似但不含所需裁决的段落。

英文摘要

Retrieval-Augmented Generation is used for Islamic question answering, but most systems are evaluated end-to-end, making retrieval failures difficult to isolate from generation failures. We study answer-bearing retrieval for Arabic fiqh, where a passage is relevant only if it states the ruling required by the question. We build a retrieval test collection for Arabic fiqh and use it to evaluate dense, lexical, hybrid, fine-tuned, and madhhab-aware retrieval strategies. The best retriever achieves 0.524 MRR@5, while fine-tuning improves performance to 0.553. Hybrid retrieval provides limited gains for strong models, whereas madhhab-aware filtering more than doubles MRR@5 on school-specific questions. We further present an error analysis showing that the main challenge is distinguishing answer-bearing passages from topically similar passages that do not contain the requested ruling.

↑