arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.15624cs.HCcs.AI

超越AI素养:胜任生成式AI使用评估工具的结构化综述与探索性元分析

Beyond AI Literacy: A Structured Review and Exploratory Meta-Analysis of Measures for Competent Generative-AI Use

Daniele Veri'

AI总结:

本研究通过结构化综述和探索性元分析,系统梳理了胜任生成式AI使用的四类评估工具,发现现有测量缺乏验证,并提出一个四层工作场所评估方案作为初步框架。

AI中文摘要:

研究人员在评估工作中的胜任生成式AI使用时,必须在自我报告、客观测试以及监督和依赖的测量之间做出选择。我们开展了一项结构化、有种子文献的综述,涵盖24篇重点实证出版物,起始于2024年基于COSMIN的综述,并补充了截至2026年8月17日的定向更新。我们将这些测量工具分为四个领域:知识与使用、认知监督、依赖校准,以及工具使用智能体的操作控制。在探索性元分析中,我们汇总了来自一个研究项目的三个直接主观-客观相关性(REML r=.055;Hartung-Knapp 95%置信区间[-.047,.156];合并报告样本量N=2,765)。我们无法解决最大研究报告的相关性与p值之间的差异,导致其权重不确定。加入第四项研究中12个跨因子相关性的合成均值后,得到r=.079(95%置信区间[-.025,.181])。此敏感性分析涉及更广泛的比较。基于这一小规模证据基础,我们无法确定总体相关性、验证工作场所的临界值,或证明用自我评分替代绩效分数的合理性。我们识别了基础知识测试(AICOS-S和GLAT)以及验证、依赖、信任和依赖性测量。在重点语料中,我们未发现经过验证的个体层面工具能够测试智能体范围、权限、恢复、状态隔离、独立审查和基于证据的关闭的完整组合;部分工具覆盖了子集。我们提出了一个具有非补偿性决策规则的四层工作场所评估方案,但尚未测试其阈值或验证其是否优于其他评估方法。

英文摘要:

Researchers assessing competent generative-AI use at work must choose among self-reports, objective tests, and measures of oversight and reliance. We conducted a structured, seeded review of 24 focal empirical publications, starting from the 2024 COSMIN-based review and adding a targeted update through 17 August 2026. We grouped the measures into four domains: knowledge and use, epistemic oversight, reliance calibration, and operational control of tool-using agents. In an exploratory meta-analysis, we pooled three direct subjective-objective correlations from one research program (REML r = .055; Hartung-Knapp 95% CI [-.047, .156]; combined reported N = 2,765). We could not resolve a discrepancy between the largest study's reported correlation and p-value, leaving its weight uncertain. Adding a synthetic mean of 12 cross-factor correlations from a fourth study gave r = .079 (95% CI [-.025, .181]). This sensitivity analysis concerns a broader comparison. From this small evidence base, we cannot establish a population correlation, validate workplace cutoffs, or justify substituting self-ratings for performance scores. We identified tests of foundation knowledge (AICOS-S and GLAT) and measures of verification, reliance, trust, and dependency. We found no validated individual-level instrument in the focal corpus that tests the full combination of agent scope, permissions, recovery, state isolation, independent review, and evidence-based closure; some cover subsets. We propose a four-layer workplace battery with non-compensatory decision rules, but have not tested its thresholds or whether it improves on other assessment approaches.

补充信息

↑