AI 中文总结
针对巴西葡萄牙语IR领域缺乏公开数据集的问题,该研究构建了NormasTCU数据集,评估了LLM作为相关性评估评判的效果,发现其在排名感知指标上与人类评判的系统排名相关性较强,可用于该场景的可扩展相关性评估。
AI 中文摘要
葡萄牙语信息检索(IR)缺乏公开数据集,且专业集合的相关性评估成本高昂。尽管大型语言模型(LLM)越来越多地用于相关性评估,但它们在非英语专业领域的可靠性仍不明确。我们推出NormasTCU(此https URL),这是一个巴西葡萄牙语IR数据集,包含14469份法律文件、46个查询,以及针对812个查询-文档对的3048个人类评判。利用NormasTCU,我们通过用两种提示技术提示三个模型来对这些对进行评分,从而评估了LLM作为相关性评估评判的效果。随后,我们比较了从LLM生成的和人类参考qrels得出的15个IR系统的排名。LLM始终表现出正向评分偏差(在0-2分制下,平均绝对误差为0.46至0.66)。此外,与人类评判的成对一致性仅达到一般至中等水平,Cohen's kappa范围为0.32至0.53。尽管存在这种偏差,LLM生成的评判通常能为nDCG@10和MRR产生高度相似的系统排名(观测到的Kendall's tau大于或等于0.90,尽管自举置信区间并非始终高于该阈值),但它们对P@10和R@10的可靠性较低。值得注意的是,基于LLM的排名有时与参考排名的相关性比单个人类注释更强。作为实际意义,我们的结果表明,当使用nDCG或MRR(排名感知指标)进行评估时,LLM可有效支持专业葡萄牙语语料库中的可扩展相关性评估,但在依赖精度或召回率时应避免使用。
英文摘要
Portuguese Information Retrieval (IR) lacks public datasets, and relevance assessment for specialized collections remains costly. While Large Language Models (LLMs) increasingly support relevance assessment, their reliability in non-English specialized domains remains unclear. We introduce NormasTCU (https://huggingface.co/datasets/LeandroRibeiro/NormasTCU), a Brazilian Portuguese IR dataset with 14,469 legal documents, 46 queries, and 3,048 human judgments over 812 query-document pairs. Using NormasTCU, we evaluated LLM-as-a-judge for relevance assessment by prompting three models with two prompt techniques to grade these pairs. We then compared the rankings of 15 IR systems derived from LLM-generated and human reference qrels. LLMs consistently showed a positive scoring bias (mean absolute error: 0.46--0.66 on a 0-2 scale). Furthermore, pair-level agreement with human judgments achieved only fair to moderate levels, with Cohen's kappa ranging from 0.32 to 0.53. Despite this bias, LLM-generated judgments often yielded highly similar system rankings for nDCG@10 and MRR (observed Kendall's tau greater than or equal 0.90, although the bootstrap confidence intervals did not always remain above this threshold), but were less reliable for P@10 and R@10. Notably, LLM-based rankings were sometimes more strongly correlated with the reference ranking than individual human annotations were. As a practical implication, our results suggest that LLMs could effectively support scalable relevance assessment in specialized Portuguese corpora when evaluated using nDCG or MRR (rank-aware metrics), but they should be avoided when relying on precision or recall.