arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.25623cs.AIcs.CYcs.HC

利用认知能力画像评估AI对职场任务的适用性

Using profiles of cognitive capability to assess AI suitability for workplace tasks

  • University of Cambridge(剑桥大学)
  • Department for Science, Innovation and Technology(科学、创新与技术部)
  • Universitat Politècnica de València(瓦伦西亚理工大学)

机构由 AI 辅助整理,请以论文原文为准。

Jonathan Prunty, Marko Tešić, Patrick Quinn, José Hernández-Orallo, Lucy Cheke

AI总结:

本研究提出基于共享核心认知能力的画像流程,通过AI与职场任务的认知维度匹配,为AI职场任务范围界定提供工具,还探讨了扩展至人机协同画像的方向。

AI中文摘要:

部署AI的组织面临范围界定问题:哪些任务可自动化、哪些应留给人类、哪些最适合两者协作。聚合基准分数几乎无法说明系统在实践中会在何处成功或失败,而人类对模型能力的判断很快就会过时。我们引入了一种流程,使用一组共享的核心认知能力对智能体和任务进行画像。认知能力画像通过对为每个项目的认知需求标注的基准测试套件的表现来推断智能体的能力;任务需求加权则从领域专家那里获取这些相同能力对其工作的相对重要性。由于两者使用相同的认知维度集合,因此可以随着模型和角色的变化独立更新,并结合起来估算AI在领域、组织、角色或单个职责层面的适用性。我们在合成智能体上验证了能力恢复情况,对6个AI系统进行了画像,并从6个职业领域的410名员工那里获取了任务需求。AI系统在认知维度上的差异大于在模型家族上的差异,而职场活动则汇聚成一个共享的认知核心。由此产生的分数提供了一种比较范围界定工具,用于识别有前景的试点候选对象以及当前系统不太可能适用的领域。我们讨论了将该框架扩展到同时对人类工人和AI系统进行画像,从AI适用性转向人机任务分配。

英文摘要:

Organisations deploying AI face a scoping problem: which tasks can be automated, which should remain with humans, and which are best shared between the two. Aggregate benchmark scores provide little insight into where systems will succeed or fail in practice, while human judgements of model capabilities quickly become outdated. We introduce a pipeline that profiles agents and tasks using a shared set of core cognitive capabilities. Cognitive capability profiling infers an agent's capabilities from performance on a benchmark battery annotated for the cognitive demands of each item. Task requirements weighting elicits from domain experts the relative importance of these same capabilities for their work. As both use a common set of cognitive dimensions, they can be updated independently as models and roles change, and combined to estimate AI suitability at the level of a domain, organisation, role, or individual duty. We validate capability recovery on synthetic agents, profile six AI systems, and elicit task requirements from 410 employees across six occupational domains. AI systems differed more across cognitive dimensions than across model families, while workplace activities converged on a shared cognitive core. The resulting scores provide a comparative scoping tool for identifying promising candidates for piloting and areas where current systems are unlikely to be well suited. We discuss extending the framework to profile human workers alongside AI systems, moving from AI suitability towards human-machine task allocation.

↑