语言模型在神经发育障碍评估中再现了人类的简化主义偏差与决策不一致性
Language Models Reproduce Human Reductionist Bias and Decision Inconsistency in Neurodevelopmental Disorders Assessment
AI总结:
研究通过对比人类与7个LLMs,发现LLMs虽表现出更高智力谦逊,却再现了人类在神经发育障碍评估中的简化主义偏差与决策不一致性,强调需审查AI的神经多样性操作化概念。
AI中文摘要:
大型语言模型(LLMs)正越来越多地为复杂的心理健康决策提供支持,这类决策不仅依赖事实证据,还依赖充满价值取向的解读。我们提出一种混合方法的人类-LLM审计框架,用于检验决策一致性、对认知启发式的敏感性、陈述性智力谦逊,以及支持分配判断所操作化的神经发育障碍概念。我们将35名人类(18名医生、17名心理学家)与7个LLMs进行对比,发现两组中对患者功能水平的评分与支持资格决策均无显著关联,这表明描述性评估与最终评价判断之间存在不一致。具体而言,两组均未表现出对锚定效应和代表性启发式的实验操纵的显著敏感性。LLMs报告的智力谦逊程度高于专家(U=241,p<.001,r=.62;LLMs:M=41.43,SD=1.99;专家:M=29.03,SD=8.05),但这与决策一致性或功能评估无关。LLMs和医生比心理学家更少授予支持(U=180.50,p=.003,r=.34),且对“基本生活需求”概念的解读不同,主要将其理解为生物生存与自我照料,而非沟通与社交需求。这些发现表明,尽管LLMs表现出高水平的智力谦逊,却再现了医学决策中嵌入的简化主义解读框架与知识。更广泛地说,我们认为在高风险情境中评估AI不仅需要测量准确性、一致性或对认知偏差的抵抗力,还需要批判性审查AI系统所操作化的神经多样性概念。
英文摘要:
Large language models (LLMs) are increasingly supporting complex mental-health decisions, which depend not only on factual evidence but also value-laden interpretations. We introduce a mixed-methods human-LLM auditing framework examining decision consistency, susceptibility to cognitive heuristics, declarative intellectual humility, and the concepts operationalized in support-allocation judgments of neurodevelopmental disorders. Comparing 35 humans (18 physicians and 17 psychologists) with seven LLMs, we show that in both groups, ratings of patients' functional level were not significantly associated with support-eligibility decisions, indicating an inconsistency between descriptive assessments and final evaluative judgments. Specifically, we find that neither group showed significant susceptibility to experimental manipulations targeting anchoring and representativeness heuristics. LLMs reported higher intellectual humility than experts (U = 241, p < .001, r = .62; LLMs: M = 41.43, SD = 1.99; experts: M = 29.03, SD = 8.05), but it was unrelated to decision consistency or functional assessment. While LLMs and physicians granted support less frequently than psychologists (U = 180.50, p = .003, r = .34), they also interpreted a concept of "basic life needs" differently, primarily as biological survival and self-care, and not communicative and social needs. These findings suggest that despite expressing high levels of intellectual humility, LLMs reproduce a reductionist interpretive framework and knowledge embedded in medical decision-making. More broadly, we argue that evaluating AI in high-stakes contexts requires not only measuring accuracy, agreement, or resistance to cognitive bias, but also critical examination of the concepts of neurodiversity that AI systems operationalize.