arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

科学智能体:评估面向特定职业的系统提示在科学任务中的表现

Scientific Agents: Evaluating Profession-Specific System Prompts on Scientific Tasks

Timothy Kassis

arXiv 2610.00084首次发表:更新:

发表机构

K-Dense(K-Dense)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究评估职业特定系统提示对科学任务的影响,发现其未提升准确性且增加成本,但长提示在API中断时更具鲁棒性。

AI 中文摘要

详细的职业特定系统提示会增加令牌使用量和每次响应的估计成本,但并未带来一致的准确性提升。我们评估了Scientific Agents,一个包含503个职业特定此HTTP URL配置文件的开放源码语料库,通过Pi智能体框架中的OpenRouter使用Gemini 3.8 Flash进行测试。我们将匹配的配置文件与四种对照进行比较:一个最小基线(“你是一个有用的助手”)、配置文件的开头角色句、一份通用的科学严谨性指南,以及一个来自不相关领域的配置文件。在九个基于文本的科学基准测试(4,531个抽样问题,100个匹配配置文件)中,4,488个项目在API错误重试后完成了所有五种条件,并使用自动化、基于规则的评分。配置文件与基线之间的平均准确率差异为-0.6个百分点(固定任务上的95%自助法置信区间为[-1.5, +0.2]),且没有任何基准显示出统计学上明确的改进。匹配的配置文件产生的输出令牌数量是基线的1.5至2.3倍,每次成功调用的成本是基线的2.2至4.5倍。在60个使用工具的BioMysteryBench生物信息学问题(基线和配置文件各运行三次)上,配置文件的平均解决率为46.7%,基线为56.7%,差异为-10.0个百分点(95%置信区间为[-16.7, -3.3]),这主要是由于配置文件下更频繁的令牌和时间限制停止所致。较长的提示有一个意外的操作优势:在SuperGPQA上,频繁的提供商API中断导致短基线仅在54.0%的项目上首次回答正确,而配置文件为71.6%。通用和不匹配的提示同样可靠,因此这一优势来自提示长度或格式,而非领域专业知识。对于所测试的模型和任务,默认加载完整的职业配置文件并不能提高准确性,且成本显著更高;选择性检索配置文件部分或开放式科学任务是否会改变这一结论仍有待测试。

英文摘要

Detailed profession-specific system prompts raise token use and estimated cost per response without a consistent accuracy gain. We evaluate Scientific Agents, an open-source corpus of 503 profession-specific AGENTS.md profiles, with Gemini 3.8 Flash via OpenRouter in the Pi agent harness. We compare matched profiles with four controls: a minimal baseline ("You are a helpful assistant"), the profile's opening role sentence, a generic scientific rigor guide, and a profile from an unrelated domain. Across nine text-based science benchmarks (4,531 sampled questions, 100 matched profiles), 4,488 items completed all five conditions after API-error retries, scored with automated, rule-based grading. The average profile-baseline accuracy difference is -0.6 percentage points (95% bootstrap interval [-1.5, +0.2] across fixed tasks), and no benchmark shows a statistically clear improvement. Matched profiles produced 1.5-2.3 times as many output tokens and cost 2.2-4.5 times more per successful call. On 60 tool-using BioMysteryBench bioinformatics problems (three runs each for baseline and profile), mean solve rates were 46.7% with the profile and 56.7% at baseline, a difference of -10.0 percentage points (95% interval [-16.7, -3.3]) driven by more frequent token- and time-limit stops under the profile. Longer prompts had one unexpected operational advantage: on SuperGPQA, frequent provider API drops left the short baseline with a correct first-pass answer on only 54.0% of items, against 71.6% with the profile. Generic and mismatched prompts were about as reliable, so this gain comes from prompt length or formatting rather than domain expertise. For the tested model and tasks, loading full profession profiles by default does not improve accuracy and costs considerably more; whether selective retrieval of profile sections or open-ended scientific tasks would change this remains to be tested.

Comments46 pages (11 pages main text, references, 33-page appendix); 10 figures, 29 tables. Evaluated corpus: https://github.com/K-Dense-AI/scientific-agents (commit 48dedd2); evaluation code and item-level records are not released

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑