arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于本体的情境化AI评估(OB-CAIE)方法论

Ontology-Based Contextual AI Evaluations (OB-CAIE) Methodology

Julie Krugler Hollek, Michael Zargham, Mala Kumar

arXiv 2610.00529首次发表:更新:

发表机构

Humane Intelligence; Dynamical Systems Group; Meta; Reliabl(Humane Intelligence; 动力系统小组; Meta; Reliabl)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

OB-CAIE方法论通过定义领域特定本体和评估过程本体,明确测试内容,平衡人类与自动化,提升AI评估的科学严谨性、可复现性及失败点可追溯性。

AI 中文摘要

基于本体的情境化AI评估(OB-CAIE)方法论旨在解决因测试覆盖不明确而导致的科学严谨性缺失问题,平衡人类专业知识与自动化,并解决AI评估测试环境缺乏可复现性的问题。OB-CAIE通过明确界定将要测试的内容,强化了当前AI评估的状态,从而解决了科学方法中的第一步。两个本体代表了OB-CAIE方法论中可处理的问题空间:领域特定本体(DSO)和评估过程本体(EPO)。DSO是“什么”(测试对象);EPO是“如何”(测试方式)。一个OB-CAIE问题空间可用于一次或多次AI评估。OB-CAIE方法论允许在特定节点、以科学依据的方式,以及在人类反馈确实不可简化或机器无法替代的复杂学科领域中,引入人类判断。OB-CAIE方法论的一个关键优势在于,失败点可以在规范的OB-CAIE方法论问题空间内被追踪、可视化和分析。

英文摘要

The ontology-based contextual AI evaluation (OB-CAIE) methodology was developed to address a lack of scientific rigor that arises from unclear testing coverage, to balance human expertise and automations, and to address a lack of reproducibility of AI evaluation testing environments. OB-CAIE strengthens the current state of AI evaluations by addressing the first step in the scientific method by clearly defining what will be tested. Two ontologies represent the tractable problem space in the OB-CAIE methodology: the Domain-Specific Ontology (DSO) and the Evaluation Process Ontology (EPO). The DSO is the what; the EPO is the how. An OB-CAIE problem space can be used for one or multiple AI evaluations. The OB-CAIE methodology allows for human judgment at specific points, in scientifically grounded ways, and in complex subject areas where human feedback is genuinely irreducible or machine irreplaceable. A key advantage of the OB-CAIE methodology is that failure points can be traced, visualized and analyzed within the canonical OB-CAIE methodology problem space.

Comments17 pages, 2 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑