发表机构
Shiraz University; Utrecht University; Wageningen University & Research(设拉子大学; 乌得勒支大学; 瓦赫宁根大学及研究中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
HugSelect是用于基础模型选择的可解释多标准决策支持框架,构建含71274个模型的知识库,经多维度评估,推荐质量与商业系统相当,推理可追溯且实用直观。
AI 中文摘要
基础模型正日益作为软件组件被重复使用,这使得模型选择成为一项关键的软件工程决策。当前的模型中心主要通过流行度指标支持模型发现,往往忽略了功能能力、操作约束以及社区感知的质量。我们认为,基础模型选择应被视为一项明确的、可审计的软件组件选择任务,而非关键词搜索、流行度排名或不透明的对话式建议。本文提出了HugSelect,一种用于基础模型选择的可解释决策支持框架。HugSelect通过将仓库元数据、提取的功能能力以及从社区讨论中得出的感知质量属性结合到统一流程中,构建了包含71274个模型的知识库。它使用加权加法模型对候选模型进行排名,该模型会公开标准级别的分数分解。我们通过流程验证、与四个基于商业大语言模型(LLM)的推荐系统的对比案例研究(44个场景)、细粒度消融实验以及探索性用户研究(样本量n=10)对HugSelect进行了评估。提取流程在功能特征上的F1分数为0.801,在质量属性映射上的准确率为0.84。HugSelect实现了模型级别的Coverage@10为0.61,家族级别的Coverage@10为0.91,表明其推荐质量与所评估的商业系统相当,且在排名质量上无显著整体差异,同时提供稳定、可追踪且可检查的推理过程。消融实验证实功能特征是检索准确率的主要驱动因素,初步用户反馈表明该框架实用且直观。
英文摘要
Foundation models are increasingly reused as software components, making model selection a critical software-engineering decision. Current model hubs primarily support discovery through popularity metrics, often neglecting functional capabilities, operational constraints, and community-perceived quality. We argue that foundation-model selection should be treated as an explicit, auditable software-component selection task rather than as keyword search, popularity ranking, or opaque conversational advice. This paper proposes HugSelect, an explainable decision-support framework for foundation-model selection. HugSelect builds a knowledge base of 71,274 models by combining repository metadata, extracted functional capabilities, and perceived quality attributes derived from community discussions into a unified pipeline. It ranks candidate models using a weighted additive model that exposes criterion-level score decompositions. We evaluated HugSelect through pipeline validation, comparative case studies against four commercial LLM-based recommendation systems (44 scenarios), fine-grained ablation, and an exploratory user study (n = 10). Extraction pipelines achieved an F1 score of 0.801 for functional features and an accuracy of 0.84 for quality-attribute mapping. HugSelect achieved a model-level Coverage@10 of 0.61 and family-level Coverage@10 of 0.91, showing recommendation quality comparable to that of the evaluated commercial systems, with no significant overall differences in ranking quality, while providing stable, traceable, and inspectable reasoning. Ablation confirmed that functional features were the main driver of retrieval accuracy, and preliminary user feedback suggests that the framework is useful and intuitive.
CommentsPages: 29, Figures: 5. Corresponding authors: Alireza Joonbakhsh (alireza.joonbakhsh@hafez.shirazu.ac.ir) and Siamak Farshidi (siamak.farshidi@wur.nl)