arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

RAIL:一种人工智能就绪水平的自动分类器

RAIL: An Automatic Classifier of the Artificial Intelligence Readiness Level

Juan Irving Vasquez, Juan Terven, Laura-Ivoone Garay-Jimenez

arXiv 2608.13428首次发表:更新:

发表机构

CIETEC-IPN; Instituto Politécnico Nacional; CICATA-QRO; UPIITA(CIETEC-IPN; 墨西哥国立理工学院; CICATA-QRO; UPIITA)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出RAIL分类器,将三类AI就绪框架统一为AIRL量表,通过大语言模型智能体小组裁决实现自动评估,测试显示其一致性好且避免高估。

AI 中文摘要

评估人工智能技术的成熟度对投资决策、项目管理和政策监测至关重要,但现有的就绪框架各不相同,难以自动应用:将技术就绪水平适配到AI时缺乏AI特有的准入标准,机器学习技术就绪水平预设了对内部流程工件的访问权限,AI/数据就绪维度模型采用的量表难以直接比较。本文做出两项贡献:其一,我们将这三个框架统一为统一人工智能就绪水平(AIRL),这是一个基于环境证据阶梯的九级有序量表,辅以维度上限(涵盖规范、数据存在性、数据质量、数据合法性、专家知识和算法成熟度),以及通用性锚定规则和明确的分配规则,从而仅通过工作的自然语言描述就能确定就绪水平。其二,我们提出RAIL(Readiness Assessment via Independent LLM-experts,即通过独立大语言模型专家进行就绪评估),这是一种将该量表落地的专家小组分类器:一名证据智能体和六名独立维度智能体,每个智能体都是具有明确范围任务的大语言模型,它们给出裁决后,由确定性最小规则进行聚合,再由首席专家在非对称权威下进行审核,可确认或降低小组的建议,但绝不会将其提升至上限以上。该方法在多项研究工作的分析中进行了测试,表现出一致性,且避免了整体式大语言模型分类器的高估问题。

英文摘要

Assessing the maturity of artificial intelligence technologies is essential for investment decisions, project management, and policy monitoring, yet the available readiness frameworks are heterogeneous and difficult to apply automatically: the adaptation of Technology Readiness Levels to AI lacks AI-specific gating criteria, the Machine Learning Technology Readiness Levels presuppose access to internal process artifacts, and AI/data readiness dimension models employ scales that resist direct comparison. This paper makes two contributions. First, we unify these three frameworks into the Unified AI Readiness Level (AIRL), a nine-level ordinal scale built on an environmental evidence ladder and complemented by dimensional caps (covering specification, data existence, data quality, data legality, expert knowledge, and algorithmic maturity) together with a generality-anchoring rule and explicit assignment disciplines, so that a readiness level becomes decidable from a natural-language description of the work alone. Second, we propose RAIL (Readiness Assessment via Independent LLM-experts), a panel-of-experts classifier that operationalizes the scale: one evidence agent and six independent dimension agents, each a large language model with a narrowly scoped mandate, deliver verdicts that a deterministic minimum rule aggregates and a chief expert reviews under asymmetric authority, confirming or lowering the panel's recommendation but never raising it above the caps. The method was tested in the analysis of several research works showing consistency and avoiding overestimation from monolithic LLM classifiers.

CommentsUnder review at journal

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑