arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SHELF:用于多任务文献计量基准测试的合成工具

SHELF: A Synthetic Harness for Multi-Task Bibliographic Benchmarking

Michael J. Bommarito

arXiv 2609.03047首次发表:更新:

AI 中文总结

SHELF是一款Python系统,可生成含62899份文档的多任务文献计量基准数据,对比多种方法表现,其模型排名可迁移,相关代码与数据已开源。

AI 中文摘要

图书馆和档案馆在人员与计算预算有限的情况下管理大量馆藏,但现有通用基准并未系统测试其文献计量相关工作,它们需要明确哪些方法适用于自身任务,以及这些方法运行所需的条件。SHELF(即 Synthetic Harness for Evaluating LLM Fitness,用于评估大语言模型适用性的合成工具)正是为填补这一空白而开发的。它是一个Python系统,可将标注分类体系、撰写规范及生成预算转化为可控的基准数据与评估任务。该首个版本包含62899份由模型生成的文档,这些文档基于美国国会图书馆词汇构建,涵盖分类、聚类、检索、成对分类及指令检索等任务。我们对比了TF、TF-IDF、BM25、主流编码器,以及仅在主题分类任务中使用的零样本解码器,每种方法仅应用于支持其的任务。实验结果显示,主题分类的准确率达0.8887,而体裁-形式分类仅为0.2605,若干成对分类与聚类任务的表现接近随机水平;稀疏方法在分类任务中仍具竞争力,TF-IDF是主题分类计时实验中测得的最快工具。此外,SHELF可独立调整文献计量维度,并能在模型训练截止日期后生成全新、可验证的未见过的文档。与LCSHBench及古腾堡项目的对比表明,模型排名的可迁移性优于绝对分数,但SHELF的分数无法用于估算生产目录数据的准确率。我们已在GitHub与Hugging Face上以宽松许可证发布了全部源代码与数据。

英文摘要

Libraries and archives manage large collections with limited staff and computing budgets, yet common benchmarks do not systematically test their bibliographic work. They need to know which methods work for their tasks and what those methods require to run. SHELF, the Synthetic Harness for Evaluating LLM Fitness, addresses this gap. It is a Python system that turns labelled taxonomies, writing specifications, and a generation budget into controlled benchmark data and evaluation tasks. This first release contains 62,899 model-written documents based on Library of Congress vocabularies, with tasks for classification, clustering, retrieval, pair classification, and instruction retrieval. We compare TF, TF-IDF, BM25, popular encoders, and, on subject classification only, zero-shot decoders; each method appears only on tasks that support it. Subject classification reaches 0.8887, while genre-form classification reaches only 0.2605, and several pair and clustering tasks remain near chance. Sparse methods remain competitive on classification, while TF-IDF is the fastest measured arm in the subject timing experiment. SHELF also varies bibliographic facets independently and can generate new, verifiably unseen documents after a model's training cutoff. Comparisons with LCSHBench and Project Gutenberg show that model rankings transfer more reliably than absolute scores, but SHELF scores do not estimate accuracy on production catalogue data. We release all source code and data under permissive licenses on GitHub and Hugging Face.

Comments17 pages, 15 tables. Code available at https://github.com/mjbommar/shelf-benchmark ; data available at https://huggingface.co/datasets/mjbommar/SHELF

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑