arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

大型语言模型在软件工程领域生成分类体系的表现如何?一个多视角评估框架

How Well Do LLMs Generate Taxonomies in the SE Domain? A Multi-perspective Evaluation Framework

Sota Nakashima, Yuta Ishimoto, Masanari Kondo, Tao Xiao, Yasutaka Kamei

arXiv 2608.01592首次发表:更新:

AI 中文总结

本文提出首个SE领域自动化分类体系生成的多视角评估框架,对比TnT-LLM与CLIMB两种方法及5种LLM,揭示二者在质量、效率上的权衡,为SE领域应用该技术提供了可操作见解。

AI 中文摘要

分类体系为软件工程(SE)研究中组织异构观测结果提供了共享的概念框架。手动构建此类分类体系耗时费力,且需要具备SE领域专业知识的标注人员。尽管大型语言模型(LLMs)的进展催生了SE领域之外的自动化分类体系生成方法,但其对技术复杂的SE工件的适用性仍不明确。在本文这篇经验性论文中,我们提出首个针对SE领域最先进自动化方法表现的全面实证评估,采用包含分类体系质量、与人类专家定义的分类体系的对齐度、独立标注下的可靠性以及效率的多视角评估框架。为支撑该评估,我们系统收集了7篇带有公开工件和人类定义分类体系的SE论文,并使用两种自动化方法(TnT-LLM和CLIMB)及5种最先进LLMs开展实验。我们的评估揭示了明确的权衡:TnT-LLM构建的分类体系质量与人类定义的相当,但成本和运行时间显著更高,且倾向于生成过于复杂的分类体系;而CLIMB速度快15--40倍、成本低8--49倍,但在需要表面相似性之外的技术推理时,质量得分往往较低。这些发现表明,TnT-LLM和CLIMB可在SE领域的实际场景中使用,而研究人员应先使用目标数据的子集评估生成分类体系的复杂性及其成本,再决定采用自动化方法还是人类专家。本研究是系统理解SE领域自动化分类体系生成的第一步,为未来研究与实践提供了可操作的见解。

英文摘要

Taxonomies provide a shared conceptual framework for organizing heterogeneous observations in software engineering (SE) research. Manually constructing such taxonomies is labor-intensive and requires annotators with expertise in the SE domain. While advances in Large Language Models (LLMs) have led to the emergence of automated taxonomy generation methods outside the SE domain, their applicability to technically complex SE artifacts remains unclear. In this experience paper, we present the first comprehensive empirical evaluation of how state-of-the-art automated methods perform on SE artifacts through a multi-perspective evaluation framework, including taxonomy quality, alignment with taxonomies defined by human experts, reliability under independent annotation, and efficiency. To support this evaluation, we systematically collect seven SE papers with publicly available artifacts and human-defined taxonomies, and conduct experiments using two automated methods (TnT-LLM and CLIMB) with five state-of-the-art LLMs. Our evaluation reveals a clear trade-off: TnT-LLM constructs high-quality taxonomies comparable to human-defined ones but incurs substantially higher cost and runtime and tends to generate overly complex taxonomies, whereas CLIMB is 15--40$\times$ faster and 8--49$\times$ cheaper but tends to score lower on quality when technical inference beyond surface-level similarity is required. These findings suggest that TnT-LLM and CLIMB can be used in practical situations in the SE domain, while researchers should first assess the complexity of the generated taxonomies and their cost using a subset of the target data to decide whether to use automated methods or human experts. Our work represents a first step toward a systematic understanding of automated taxonomy generation in SE, offering actionable insights for future research and practice.

Comments13 pages, Accepted at the 41st IEEE/ACM International Conference on Automated Software Engineering (ASE 2026), Research Papers Track

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑