发表机构
Institute for Language and Speech Processing, Athena Research Center; National Technical University of Athens(雅典娜研究中心语言与言语处理研究所; 雅典国立技术大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究引入基于布鲁姆教育目标分类学的推理步骤自动标注框架,对LRMs开展大规模分析,发现思维类型与推理正确性相关,为提升推理质量提供洞见。
AI 中文摘要
大型推理模型(LRMs)彻底变革了大型语言模型(LLMs)的推理能力,且日益公开可用的推理轨迹为研究模型行为创造了宝贵机会,不仅可从表面层面,还能深入到单个推理步骤的粒度。然而,理解推理过程中运用的思维类型——这能为模型的推理模式提供关键洞见并实现可操作的应用——仍未得到充分探索。为解决这一差距,我们引入了一个通过布鲁姆教育目标分类学视角自动标注推理步骤的框架,该框架将思维分为六个认知层级,如记忆(Remembering)、应用(Applying)和评估(Evaluating)。利用该框架,我们在多个模型和数据集上开展了大规模分析,揭示了不同模型和任务间思维模式的异同。此外,我们证明了从推理轨迹中提取的思维类型信息与推理正确性相关联,为改进推理铺平了道路。我们的研究成果为分析LRMs的思维模式建立了细粒度框架,并为提升推理质量提供了可操作的洞见。
英文摘要
Large Reasoning Models (LRMs) have revolutionized reasoning in LLMs, and the increasing public availability of reasoning traces creates valuable opportunities to study model behavior not only at the surface level but also at the granularity of individual reasoning steps. However, understanding the types of thinking employed during reasoning - which offers critical insights into models' reasoning patterns and enables actionable applications - remains underexplored. To address this gap, we introduce a framework for automatic annotation of reasoning steps through the lens of Bloom's Taxonomy, which classifies thinking into six cognitive levels, such as Remembering, Applying and Evaluating. Using this framework, we perform a large-scale analysis across models and datasets, revealing both similarities and differences in thinking patterns across models and tasks. Moreover, we demonstrate that thinking-type information derived from reasoning traces correlates with correctness, paving the way for improved reasoning. Our findings establish a fine-grained framework for analyzing thinking patterns in LRMs and provide actionable insights for enhancing reasoning quality.