arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

心理能力作为人工智能评估中缺失的维度

Psychological Competence as a Missing Dimension in AI Evaluation

Marcos Economides, Paul M. Sacher, Samuel Salzer, Alexis Michelle Abellar, Fendi Tsim, Antoine Ferrère

arXiv 2607.08285首次发表:更新:

AI 中文总结

研究指出当前人工智能评估框架对面向人类的系统不足,引入心理能力这一缺失维度,借鉴多领域研究概述其概念框架,定义结构、阐明边界并说明评估方式,强调其应成为多方关注的核心考量。

AI 中文摘要

当前的人工智能评估框架主要关注技术性能,如准确性、鲁棒性、推理能力和政策合规性。但对于通过自然语言与用户直接交互的系统来说,这些措施还不够。面向人类的人工智能系统越来越多地充当顾问、教练、导师和伙伴,其响应会影响用户的推理、情绪解读、信念形成、信任校准和决策。本文引入心理能力作为人工智能评估中缺失的维度,定义其为面向人类的人工智能系统以适合用户、情境和交互目的的方式支持用户认知、情绪解读和行为决策的能力,包括框架、语气等交互属性。现有评估方法很少直接评估这些心理影响。本文借鉴行为科学和人机交互研究,概述了心理能力及其核心领域的概念框架,定义了该结构,阐明了其边界,并描述了通过基于场景的探测、结构化人工评估和模型辅助评估方法进行评估的方式。作者认为心理能力应成为模型提供者、部署组织、研究人员和监管机构的核心考量因素。

英文摘要

Current AI evaluation frameworks focus primarily on technical performance, including accuracy, robustness, reasoning ability, and policy compliance. These measures remain essential, but they are not sufficient for systems that interact directly with users through natural language. Human-facing AI systems are increasingly used as advisors, coaches, tutors, and companions. In these roles, their responses can shape how users reason, interpret emotions, form beliefs, calibrate trust, and make decisions. The relevant unit of evaluation is therefore not only the model, but the human-AI interaction. This paper introduces psychological competence as a missing dimension in AI evaluation. We define psychological competence as the capacity of a human-facing AI system to support user cognition, emotional interpretation, and behavioral decision-making in ways that are appropriate to the user, context, and purpose of the interaction. This includes interaction properties such as framing, tone, perceived authority, responsiveness, uncertainty handling, and conversational guidance. Existing evaluation approaches capture parts of this problem but rarely assess these psychological effects directly. Drawing on behavioral science and human-AI interaction research, we outline a conceptual framework for psychological competence and its core domains. Rather than proposing a specific benchmark, we define the construct, clarify its boundaries, and describe how it may be assessed through scenario-based probes, structured human evaluation, and model-assisted evaluation methods. We argue that psychological competence should become a core consideration for model providers, deploying organizations, researchers, and regulators concerned with the real-world effects of human-facing AI systems.

Comments12 pages, 3 figures. LaTeX source added to enable accessible HTML; minor wording and typesetting revisions

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑