arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

大型语言模型是否偏向特定研究主题?

Do Large Language Models Favour Any Research Topics?

Mike Thelwall

arXiv 2609.00323首次发表:更新:

发表机构

University of Sheffield(谢菲尔德大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究以73489篇健康与生命科学期刊文章为样本,发现GPT-OSS-120B与Gemma 3 27B存在主题偏好差异,证明LLM用于研究评估时存在偏差,需重视该问题。

AI 中文摘要

大型语言模型(LLM)可评估已发表期刊文章的质量,在需要评估时可辅助人工评估。虽然有理由认为LLM在该角色中可能存在偏差,但目前尚无统计层面的有力证据。本研究针对该空白,探索来自15种健康与生命科学期刊的73489篇文章中,哪些类型的文章会获得LLM的高分或低分。通过以多种方式对比两个LLM的高分与低分文章标题及摘要中的词汇,结果显示,GPT-OSS-120B偏向的主题包括病毒、基因和细胞,其不偏向的主题包括综述、患者和学生。不过,目前尚不清楚这些模式是反映了潜在的质量差异还是AI偏差。相同方法发现GPT-OSS-120B与Gemma 3 27B偏向的主题存在系统性差异,例如Gemma 3 27B对机器学习研究给出相对更高的分数,证明两者中至少有一个存在AI偏差。最后,对比全文文章的评分与仅标题和摘要的评分,发现两个LLM均存在差异,表明它们对至少一种输入类型(全文或仅标题摘要)会表现出AI偏差,且可能两者都存在。总体而言,结果显示在决定是否将LLM用于研究评估任务时,考虑LLM的偏差十分重要。

英文摘要

Large Language Models (LLMs) can estimate the quality of published journal articles, potentially supporting human assessment when evaluations are needed. Whilst there are reasons to believe that LLMs may have biases in this role, there is no statistically strong evidence yet. The current article addresses this gap with an exploration of the types of articles that attract high or low LLM scores in 73,489 articles from 15 health and life sciences journals. Based on comparing the words in the titles and abstracts of higher and lower scoring articles for two LLMs in various ways, the results suggest that topics favoured by GPT-OSS-120B include viruses, genes and cells and its disfavoured topics include surveys, patients and students. It is not clear whether these patterns reflect underlying quality differences or AI biases, however. The same method found systematic differences between the topics favoured by GPT-OSS-120B and Gemma 3 27B, such as Gemma 3 27B giving relatively higher scores for machine learning research, proving that at least one of the two LLMs has AI bias. Finally, comparing the scores for full-text articles compared to scores for titles and abstracts also finds differences for both LLMs, showing that they both can exhibit AI bias for at least one of these two input types, and probably both. Overall, the results show that it is important to consider LLM biases when deciding whether to use them for research evaluation tasks.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑