arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.12403cs.AIcs.CL

超越ID嵌入:面向认知诊断的过程接地语言建模

Beyond ID Embeddings: Process-Grounded Language Modeling for Cognitive Diagnosis

Minghang Liu, Yuanzhuo Wang, Qiang Qiu, Huawei Shen, Xueqi Cheng

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出PLCD框架,利用语言结构作为认知先验,结合响应记录校准学生状态,通过LLM构建概念图式和过程图,并采用DA-MoE和对比学习映射到认知空间,在预测表现和认知迁移上超越传统CDMs。

中文摘要 AI 辅助

认知诊断模型(CDMs)在个性化在线学习中发挥着关键作用。传统CDMs依赖离散的、基于ID的嵌入来表示学生、习题和概念。这种范式与学习者认知的本质相背离,因为知识并非作为孤立的符号被存储和检索。因此,当新习题或概念出现时,CDMs面临语义局限性。在本文中,我们提出了一种过程感知的语言认知诊断(PLCD)框架,该框架使用语言派生的结构作为认知先验,并利用响应记录来校准学生后验状态。PLCD利用大型语言模型(LLMs)构建概念图式和认知过程图,并使用目标条件语义记忆来检索与每个目标习题相关的历史响应。随后,一个带有DA-MoE专家和过程级对比学习的过程接地语言到认知映射器将文本证据映射到统一的认知空间。实验结果表明,PLCD不仅在预测学生表现方面优于传统基线,而且展现出强大的认知迁移能力。这些结果将LLMs的计算能力与测量潜在知识状态的心理测量目标联系起来,表明由响应记录校准的结构化语言先验可以提高冷启动鲁棒性和认知接地性。

英文摘要

Cognitive Diagnosis Models (CDMs) play a pivotal role in personalized online learning. Traditional CDMs rely on discrete, ID-based embeddings to represent students, exercises, and concepts. This paradigm diverges from the nature of learner cognition, where knowledge is not stored and retrieved as isolated symbols. As a result, CDMs suffer from semantic limitations when new exercises or concepts appear. In this paper, we propose a Process-aware Language Cognitive Diagnosis (PLCD) framework that uses language-derived structures as cognitive priors and response records to calibrate student posterior states. PLCD leverages large language models (LLMs) to construct concept schemas and cognitive process graphs, and uses target-conditioned semantic memory to retrieve historical responses that are relevant to each target exercise. A process-grounded Language-to-Cognition Mapper with DA-MoE experts and process-level contrastive learning then maps the textual evidence into a unified cognitive space. Experimental results show that PLCD not only outperforms traditional baselines in predicting student performance but also exhibits strong cognitive transfer capabilities. These results connect the computational power of LLMs with the psychometric goal of measuring latent knowledge states, suggesting that structured language priors calibrated by response records can improve cold-start robustness and cognitive grounding.

发表机构

  • State Key Laboratory of AI Safety, Institute of Computing Technology, CAS(中国科学院计算技术研究所人工智能安全国家重点实验室)
  • University of Chinese Academy of Sciences(中国科学院大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑