arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

学术导师发现的检索方法比较:针对美国9所高校768名计算机科学教师档案的六种方法研究

Comparing Retrieval Methods for Academic Advisor Discovery: A Six-Method Study of 768 CS Faculty Profiles Across 9 US Universities

Biraj Subedi

arXiv 2609.03901首次发表:更新:

AI 中文总结

本研究针对学术导师发现任务,对比评估六种检索方法,基于768名美高校CS教师档案数据,发现Reranked方法表现最优,还揭示了不同信息类型对检索效果的影响并公开了全部相关资源。

AI 中文摘要

我们针对学术导师发现任务开展了六种信息检索方法的对比评估:即根据研究生申请者的研究兴趣陈述对计算机科学(CS)教师进行相关性排名。所评估的方法涵盖稀疏词汇匹配(Jaccard重叠、TF-IDF、BM25)、密集语义检索(all-MiniLM-L6-v2句子嵌入)、混合分数融合以及学习排序。评估采用了一个新的领域特定数据集:从美国9所高校计算机科学系爬取的768份教师档案,包含针对5个代表不同研究生研究概况的查询的162个分级相关性判断(等级0/1/2)。在所有5个查询中,Reranked方法获得最高的平均NDCG@10(0.477,标准差0.138),其次是Semantic(0.450)、Hybrid(0.421)、BM25(0.406)、Jaccard(0.303)和TF-IDF(0.246)。在所有15组两两比较经Bonferroni校正后,TF-IDF的表现显著差于BM25、Semantic、Hybrid和Reranked;其余两两差异在5个查询的校正后均不显著。领域 ablation分析显示,仅个人简介(NDCG 0.634)的表现优于结合个人简介与研究领域标签的完整模型(0.593)。对照实验表明,拼接arXiv论文摘要会使NDCG@10降低0.176,这为晚融合架构提供了动机。所有代码、爬虫工具及相关性标签均已公开发布。

英文摘要

We present a comparative evaluation of six information retrieval methods for the task of academic advisor discovery: ranking CS faculty members by relevance to a graduate applicant's research interest statement. The methods span sparse lexical matching (Jaccard overlap, TF-IDF, BM25), dense semantic retrieval (all-MiniLM-L6-v2 sentence embeddings), hybrid score fusion, and learning-to-rank. Evaluation uses a new domain-specific collection: 768 faculty profiles scraped from 9 US CS departments, with 162 graded relevance judgments (grade 0/1/2) across 5 queries representing distinct graduate student research profiles. Across all five queries, Reranked achieves the highest mean NDCG@10 (0.477, std 0.138), followed by Semantic (0.450), Hybrid (0.421), BM25 (0.406), Jaccard (0.303), and TF-IDF (0.246). After Bonferroni correction across all 15 pairwise comparisons, TF-IDF is significantly worse than BM25, Semantic, Hybrid, and Reranked; no other pairwise difference survives correction at 5 queries. A field ablation reveals that biography alone (NDCG 0.634) outperforms the full model combining biography with research area tags (0.593). A controlled experiment shows that concatenating arXiv paper abstracts reduces NDCG@10 by 0.176, motivating a late-fusion architecture. All code, scrapers, and relevance labels are released openly.

Comments16 pages, 2 figures, 9 tables. Code and demo at https://github.com/subedibiraj/academic-discovery

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑