发表机构
Guangdong Institute of Smart Education, Jinan University; School of Physical Education, Jinan University; Guangdong Provincial Key Laboratory of Speed Capability Research; Sapient Intelligence Pte Ltd(暨南大学广东智慧教育研究院; 暨南大学体育学院; 广东省速度能力研究重点实验室; Sapient Intelligence Pte Ltd)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出SRJudge三阶段框架,通过选择器缩小候选、推理器精细推理、评判器评估,增强大语言模型在细粒度知识概念标注中的选择性推理能力,并在生物和物理数据集上超越现有基线。
AI 中文摘要
知识概念标注旨在为教育内容分配特定的概念或主题标签,这在传统和在线教学实践中对教育者和学习者都至关重要。近期研究已探索将大语言模型(LLMs)应用于此任务,并取得了令人瞩目的性能。然而,由于决策空间的高维度,LLMs在从大规模候选集中选择正确概念时仍面临困难。在本文中,我们提出了一种新颖的三阶段“选择-推理-评判”(SRJudge)框架,该框架赋予LLMs选择性推理能力,用于细粒度知识概念标注。具体而言,第一阶段的“选择器”首先通过微调一个小语言模型(SLM),如BERT,将候选概念缩小至top-K候选列表,因为top-K预测在大多数情况下能命中正确概念,从而减少正确候选的决策空间。接下来,第二阶段的“推理器”采用轻量级LLM对候选列表进行精细推理。它进一步整合了改进的强化学习策略,包括动态任务特定奖励函数和剪枝机制,以更好地与人类推理偏好对齐。最后,一个更大的LLM充当“评判器”,评估推理过程及其解释的整体合理性,以确定最终输出。此外,我们构建了两个高质量数据集用于进一步验证,即生物数据集S_Bio和物理数据集S_Phy。实验结果表明,我们的方法在基准数据集上持续优于最先进的基线方法,验证了其有效性和优越性。资源可在以下网址获取:this https URL。
英文摘要
Knowledge concept tagging aims to assign specific concept or topic labels to educational content, which is essential for both educators and learners in traditional and online teaching practices. Recent work has explored large language models (LLMs) for this task, achieving promising performance. However, LLMs still struggle to select the correct concept from a large-scale candidate set due to the high dimensionality of the decision space. In this paper, we propose a novel three-stage Select-Reason-Judge (SRJudge) framework, which empowers LLMs with selective reasoning capability for fine-grained knowledge concept tagging. Specifically, the Selector in Stage 1 first narrows the candidate concepts to a top-K shortlist by fine-tuning a small language model (SLM), e.g., BERT, since the top-$K$ predictions hit the correct concept in most cases, thereby reducing the decision space of correct candidates. Next, the Stage 2 Reasoner employs a lightweight LLM for refined reasoning over the shortlisted candidates. It further integrates an improved reinforcement learning strategy with a dynamic task-specific reward function and a pruning mechanism to better align with human reasoning preferences. Finally, a larger LLM acts as a judger that evaluates the overall rationality of the reasoning process and its explanations to determine the final output. In addition, we construct two high-quality datasets for further validation, i.e., the biology dataset S_Bio and the physics dataset S_Phy. Experimental results demonstrate that our method consistently outperforms state-of-the-art baselines across benchmark datasets, verifying its effectiveness and superiority. Resources are available at: https://github.com/Nicozwy/SRJudge.
CommentsAccepted by IJCAI 2026