人工智能认知的多维度评估(MAAC):面向基于文本的人工智能系统的过程导向认知评估理论框架
Multi-Dimensional Assessment for AI Cognition (MAAC): A Theoretical Framework for Process-Oriented Cognitive Evaluation of Text-Based AI Systems
浏览论文内容
中文总结 AI 辅助
针对现有AI评估仅关注结果的局限,本文提出MAAC框架,从九个认知维度对基于文本的AI系统进行过程导向评估,补充了现有结果基准的不足。
中文摘要 AI 辅助
人工智能系统的评估历来依赖基于结果的基准,这些基准用于衡量任务准确率、鲁棒性或公平性。尽管这些基准不可或缺,但它们对产生性能的潜在认知过程提供的诊断见解有限,留下了关于人工智能系统如何推理、整合记忆、管理复杂性或避免生成虚假信息的关键问题未得到解答。本文介绍了人工智能认知的多维度评估(MAAC),这是一个以理论为基础的框架,用于将评估从基于文本的人工智能系统产出转向其思考过程。MAAC定义了九个受认知科学启发的维度:认知负荷、工具执行、内容质量、记忆整合、复杂性处理、幻觉控制、知识迁移、处理效率以及过程-结果对齐。每个维度都基于已确立的认知科学理论,包括马尔的三层假说、巴德利的工作记忆模型、斯威勒的认知负荷理论以及认知统一理论。五项理论分析为该框架的连贯性和经验可测试性提供了初步支持:维度-理论映射;评估广度和非冗余性的覆盖矩阵;相对于当前评估实践的正式差距分析;一个已完成的诊断示例;以及一组用于未来经验测试的先验相互依赖性预测。MAAC为基于文本的人工智能系统提供了一个原则性的过程层面认知评估的理论和操作框架,用基于认知的多维度评估补充了现有的基于结果的基准。
英文摘要
Evaluating artificial intelligence systems has historically relied on outcome-based benchmarks that measure task accuracy, robustness, or fairness. While indispensable, these benchmarks provide limited diagnostic insight into the underlying cognitive processes that generate performance-leaving critical questions unanswered about how AI systems reason, integrate memory, manage complexity, or avoid generating false information. This paper introduces the Multi-Dimensional Assessment for AI Cognition (MAAC), a theoretically grounded framework for shifting evaluation from what text-based AI systems produce to how they think. MAAC defines nine cognitively motivated dimensions: Cognitive Load, Tool Execution, Content Quality, Memory Integration, Complexity Handling, Hallucination Control, Knowledge Transfer, Processing Efficiency, and Process-Outcome Alignment. Each dimension is grounded in established cognitive science theory-drawing on Marr's tri-level hypothesis, Baddeley's working memory model, Sweller's cognitive load theory, and unified theories of cognition. Five theoretical analyses provide initial support for the framework's coherence and empirical testability: dimension-to-theory mapping; a coverage matrix assessing breadth and non-redundancy; a formal gap analysis relative to current evaluation practice; a worked diagnostic illustration; and a set of a priori interdependency predictions for future empirical testing. MAAC provides a theoretical and operational framework for principled process-level cognitive assessment of text-based AI systems, complementing existing outcome-based benchmarks with cognitively grounded, multi-dimensional evaluation.
发表机构
- Wayne State University(韦恩州立大学)
机构由 AI 辅助整理,请以论文原文为准。