科学声明-来源检索再探讨:风格迁移与重新排序的比较研究
Scientific Claim-Source Retrieval Revisited: A Comparative Study of Style Transfer and Re-Ranking
浏览论文内容
中文总结 AI 辅助
研究社交媒体科学声明源检索难题,对比稀疏和密集检索模型,发现翻译声明、纳入元数据、风格迁移及新型重新排序方法可提升检索性能,基于验证的重新排序效果最佳,MRR@5达0.758。
中文摘要 AI 辅助
社交媒体上的科学声明往往难以核实,可能助长错误信息传播。为应对这一挑战,自动事实核查系统需要科学声明-来源检索,即识别给定声明背后的源出版物。然而,声明与其源出版物在语言、风格和特异性上往往差异很大,检索具有挑战性。我们在CheckThat! 2026基准上对稀疏和密集检索模型进行了科学声明-来源检索的比较研究。结果表明,将声明翻译成英文比原始和双语声明表示更优,纳入出版物元数据可通过捕获间接源引用获得额外检索增益。此外,分析了四种风格迁移方法,发现它们对大多数模型都能提高检索性能,尽管最佳风格取决于底层检索目标。最后,研究了基于相似性和信号的重新排序方法,引入了三种基于归因、实体重叠和基于验证的推理的新型重新排序模型。基于验证的重新排序在语义相似性之外还能带来额外增益,以0.758的MRR@5实现了最佳整体性能。
英文摘要
Scientific claims shared on social media are often difficult to verify and may contribute to the spread of misinformation. To address this challenge, automated fact verification systems require scientific claim-source retrieval, the task of identifying the source publication underlying a given claim. However, claims often differ substantially from their source publications in language, style, and specificity, making retrieval challenging. We present a comparative study of scientific claim-source retrieval on the CheckThat! 2026 benchmark across sparse and dense retrieval models. Our results show that translating claims into English outperforms both original and bilingual claim representations, while incorporating publication metadata provides additional retrieval gains by capturing indirect source references. In addition, we analyze four style transfer approaches and find that they improve retrieval performance for most models, although the optimal style depends on the underlying retrieval objective. Finally, we investigate similarity- and signal-based re-ranking approaches, introducing three novel re-ranking models based on attribution, entity overlap, and verification-based reasoning. Verification-based re-ranking yields additional gains beyond semantic similarity and achieves the best overall performance with an MRR@5 of 0.758.