arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

从孤立任务到结构化能力:大语言模型的多层分类法

From Isolated Tasks to Structured Capabilities: A Multilayer Taxonomy for Large Language Models

Shixin Fang, Jiachen Wo, Wenjuan Qin, Sihang Jiang, Yanghua Xiao

arXiv 2607.22182首次发表:更新:

发表机构

Fudan University(复旦大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出大语言模型多层分类法,含14个能力领域和91个子技能,以人类认知科学为指导。通过筛选和映射相关论文,分析研究关注情况,揭示领域及子技能分布,此分类法有助于大语言模型研究组织、评估等多方面工作。

AI 中文摘要

大语言模型(LLM)评估涵盖各种任务和基准,但相关证据仍围绕任务而非所探测的能力组织。这种碎片化限制了跨研究比较,掩盖了任务所调用的能力,难以识别覆盖差距。我们引入了一个多层分类法,包含原始、构建和整合层的14个能力领域和91个子技能。人类认知科学指导能力定义和组织,而非LLM架构。通过对2023年至2025年间来自ACL、AAAI、ICML和NeurIPS的31,505篇论文进行筛选,并通过多模型注释、共识和仲裁对15,934篇以LLM为重点的论文进行映射。研究关注集中在语言语义能力、推理、规划和决策制定以及感知等方面,六个领域出现频率较低。该分类法支持研究组织、覆盖审核、评估解释以及诊断、训练和转移的可测试假设。

英文摘要

Large language model (LLM) evaluation spans diverse tasks and benchmarks, yet evidence remains organized around tasks rather than the capabilities they probe. This fragmentation limits cross-study comparison, obscures capabilities tasks recruit, and makes coverage gaps difficult to identify. We introduce a multi-layer taxonomy of 14 capability domains and 91 subskills across Primitive, Constructed, and Integrative layers. Human cognitive science guides capability definition and organization, not LLM architecture. Layer assignments draw on developmental precedence and hypothesized functional support, while human-origin constructs are adapted to observable model behavior. To demonstrate operational utility, we screened 31,505 papers from ACL, AAAI, ICML, and NeurIPS between 2023 and 2025 and mapped 15,934 LLM-focused papers through multi-model annotation, consensus, and arbitration. Direct research attention concentrated on Language-Semantic Competence (3,551; 22.3%), Reasoning (3,388; 21.3%), Planning and Decision-Making (2,149; 13.5%), and Perception (1,954; 12.3%), whereas six domains appeared in fewer than 2% of papers. Within domains, the most frequent subskill had a median prevalence of 97.9% and appeared in at least 90% of papers in 10 of 14 domains. Language-Semantic Competence and Reasoning formed the highest-volume pair (n = 1,864; 11.7%; lift = 2.47), whereas Theory of Mind and Social Reasoning and Interaction showed the highest lift among pairs with at least 20 co-occurrences (n = 62; lift = 30.84). By shifting the unit of analysis from isolated tasks to structured capabilities, the taxonomy supports research organization, coverage audits, evaluation interpretation, and testable hypotheses for diagnosis, training, and transfer.

Comments34 pages, 5 figures, 20 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑