AI 中文总结
研究无监督、基于内容的学术合作推荐,比较TF-IDF基线、主题模型和基于嵌入的检索等方法,引入受限设置评估模型,还从两视角考察可解释性,展现不同方法差异及权衡。
AI 中文摘要
本文研究学术环境中使用出版物文本的无监督、基于内容的合作推荐。比较三类方法:TF-IDF基线、基于主题的模型(LDA和BERTopic,包括克隆变体)以及使用SciBERT和Faiss的基于嵌入的检索。引入受限设置评估模型行为,通过两种视角考察可解释性,结果显示方法间存在明显差异。
英文摘要
In this paper, we examine unsupervised, content-based collaboration recommendations using publication text in scholarly settings. We compare three families of methods: a TF-IDF baseline, topic-based models (LDA and BERTopic, including clone variants), and embedding-based retrieval using SciBERT with Faiss. To evaluate model behavior beyond simple lexical matching, we introduce a constrained setting where publication overlap between researchers is partially removed while still using historical co-authorship as proxy ground truth for post-hoc evaluation. Results show clear differences across methods. TF-IDF performs best under full information but drops significantly as overlap is reduced. In contrast, topic-based and embedding-based approaches show more stable performance, suggesting they capture broader distributional similarities, rather than relying only on direct lexical overlap. We also examine explainability through two perspectives: intrinsic topic-based explanations and post-hoc, retrieval-based explanations generated using language models. These provide complementary trade-offs between transparency and human readability.
Comments6 pages, 2 figures, Submitted to ICMLA 2026