arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

希腊图书出版商环境中嵌入模型和大语言模型的比较评估 - CUP数据集

A Comparative Evaluation of Embeddings and LLMs in a Greek Book Publisher Setting - The CUP Dataset

Katerina Papantoniou, Panagiotis Papadakos, Theodore Patkos, Dimitris Garefalakis, Nikos Vardakis, Dimitris Plexousakis

arXiv 2607.21274首次发表:更新:

发表机构

ICS-FORTH; Crete University Press(希腊计算机科学与技术研究所; 克里特大学出版社)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出希腊图书检索基准CUP数据集,评估稀疏、密集、混合及大语言模型辅助的检索方法,发现多语言嵌入优于希腊特定模型,混合检索最佳,还分析了不同方法在各类查询中的表现及大语言模型相关技术的效果。

AI 中文摘要

我们展示了CUP,一个希腊图书检索基准,由868条目录记录和104个带有分级相关性判断的专家注释查询组成。我们在这个图书搜索环境中评估了稀疏(BM25)、密集(句子转换器)、混合和大语言模型辅助的检索方法。多语言嵌入优于希腊特定模型,而混合检索总体表现最佳。查询级分析表明,BM25在命名实体查询方面表现出色,而密集和混合方法改善了自然语言、嘈杂、跨语言和概念查询。字段感知提示有模型特定的效果,而大语言模型目录摘要改善了仅目录的检索,大语言模型后过滤以高成本改善了早期检索。总体而言,CUP能够对希腊语检索进行跨词汇、语义、嘈杂和跨语言查询的真实世界评估。

英文摘要

We present CUP, a Greek book retrieval benchmark consisting of 868 catalog records and 104 expert-annotated queries with graded relevance judgments. We evaluate sparse (BM25), dense (sentence-transformers), hybrid, and LLM-assisted retrieval methods in this book-search setting. Multilingual embeddings outperform Greek-specific models, while hybrid retrieval performs best overall. A query-level analysis shows that BM25 excels at named-entity queries, while dense and hybrid methods improve natural-language, noisy, cross-lingual, and concept queries. Field-aware prompting has model-specific effects, while LLM TOC summarization improves TOC-only retrieval and LLM post-filtering improves early-stage retrieval at a high cost. Overall, CUP enables real-world evaluation of Greek retrieval across lexical, semantic, noisy, and cross-lingual queries.

CommentsPreprint of a manuscript submitted to the 14th EETN Conference on Artificial Intelligence (SETN 2026)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑