评估叙述者:多智能体知识系统中声明级来源的伊纳德-里贾尔框架
Grading the Narrators: An Isnad-Rijal Framework for Claim-Level Provenance in Multi-Agent Knowledge Systems
浏览论文内容
中文总结 AI 辅助
研究多智能体知识系统声明级来源问题,借鉴经典伊斯兰圣训科学方法,贡献形式映射、关系模式、决策矩阵等,对20000条声明评估,验证部分方法,也发现等级恢复循环部分失败及两次分析无定论。
中文摘要 AI 辅助
现代多智能体知识系统通过自主转换链积累知识,而非直接检索。现有来源工作记录发生的事情,源可靠性估计已确立。缺少的是一个操作框架,为声明级传输链附加分级、按域的传输器可靠性,具有完整性语义、转换类型聚合、解耦内容批评和服务/审查/隔离路由。经典伊斯兰圣训科学面临类似问题,发展出严谨方法。本文将该方法转移到人工智能系统设计。贡献了从圣训科学概念到多智能体管道的形式映射、实现声明链和分级叙述者注册表的关系模式、结合链等级与内容批评的决策矩阵,并对真实物理教科书的20000条声明进行评估。评估验证了最弱链隔离和独立链确证;报告了等级恢复循环的部分失败及两次分析无定论。
英文摘要
Modern multi-agent knowledge systems increasingly accumulate knowledge through chains of autonomous transformations rather than direct retrieval. Existing provenance work records what happened - execution traces, tool calls, evidence links - and source-reliability estimation is long established (truth discovery, reputation systems). What is missing is an operational framework that attaches graded, per-domain transmitter reliability to claim-level transmission chains, with completeness semantics, transformation-typed aggregation, decoupled content criticism, and serve/review/quarantine routing. Classical Islamic hadith science confronted a structurally similar problem: deciding whether knowledge transmitted through chains of human narrators should be accepted. Over centuries it developed a rigorous methodology - isnad (a complete transmission chain attached to every claim), rijal (systematic grading of each narrator's integrity and precision), weakest-link chain evaluation, corroboration through independent chains, and matn criticism (content evaluated independently of chain quality). This paper transfers that methodology to AI system design. We contribute a formal mapping from hadith-science concepts to multi-agent pipelines, a relational schema implementing claim chains and a graded narrator registry, a decision matrix combining chain grade with content criticism, and an evaluation on 20,000 claims from real physics textbooks. The evaluation validates weakest-link quarantine and independent-chain corroboration; reports a partial failure of the grade-recovery loop, which missed the highest-fault narrator; and reports two analyses as inconclusive, including a matched-coverage comparison the framework could not reach with the reference content critic. The paper is explicit throughout about which claims the evidence does and does not yet support.