arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

利用四象限方法评估Redpine Science

Leveraging a four-quadrant approach for evaluating Redpine Science

Filip Dorm, Leonora Vesterbacka

arXiv 2610.07937首次发表:更新:

发表机构

Redpine(Redpine)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究采用四象限方法,通过公开和专家验证基准评估Redpine Science的检索相关性与回答质量,结果显示其在多项指标上优于无检索和网络搜索基线。

AI 中文摘要

Redpine Science为模型和智能体提供了对广泛同行评审文献的单一访问点,可通过模型上下文协议(MCP)和API直接查询。本报告从两个层面评估Redpine Science:检索块的相关性,以及模型在访问Redpine Science与使用网络搜索时的回答质量。评估使用了公开基准和专家验证基准。公开基准是测试模型开发的广泛接受方式,可在不同实验室间进行比较,但存在饱和与记忆风险。为解决这一问题,我们补充了专家验证的问题集。本报告共呈现四项评估。在公开答案质量基准ScholarQABench SciFact上,使用Redpine Science的智能体正确回答了94.4%的声明,而无检索时为87.6%。在专家验证的问题集上,使用Redpine Science的智能体陈述了80.1%的必要声明,而仅限网络搜索的智能体为70.2%。在包含668个查询的公开检索基准中(其黄金论文由Redpine持有),去除任何模型推理后,Redpine Science将正确来源论文置于前十结果中的比例为83.1%(Recall@10),而基准创建者为79.3%。一个盲法专家相关性小组评定Redpine Science的Precision@5为75.2%,而PubMed搜索工具为39.8%。我们发布了专家验证的问题集以及复现上述所有主要结果的说明,网址为https URL。

英文摘要

Redpine Science gives models and agents a single access point to a wide range of peer-reviewed literature, queried directly through the Model Context Protocol (MCP) and an API. This report evaluates Redpine Science on two levels: the relevance of the retrieved chunks, and a model's answer when it has access to Redpine Science compared to web search. Both public and expert-validated benchmarks are used. Public benchmarks are a widely accepted way to test model development and are comparable across labs, but risk saturation and memorization. To address this, we complement them with an expert-validated question set. In total, this report presents four evaluations. On ScholarQABench SciFact, the public answer-quality benchmark reported here, an agent with Redpine Science answers 94.4% of claims correctly against 87.6% with no retrieval. On the expert-validated question set, an agent with Redpine Science states 80.1% of the required claims against 70.2% for an agent restricted to web search. On the 668 queries of a public retrieval benchmark whose gold paper Redpine holds, stripped of any model reasoning, Redpine Science places the correct source paper in its top ten results for 83.1% of queries (Recall@10), against 79.3% for the benchmark's creator. A blinded expert relevance panel places Redpine Science's Precision@5 at 75.2% against 39.8% for the PubMed search tool. We release the expert-validated question set and instructions to reproduce every headline result above, at https://github.com/redpine-ai/benchmarks.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑