K-Bench:衡量模型在真实科学智能体请求上的性能
K-Bench: measuring model performance on real scientific agent requests
查看机构详情
- K-Dense, Inc.(K-Dense公司)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
本研究构建K-Bench 01基准,评估9个前沿模型处理真实科学请求的性能,发现无模型达领域科学家接受阈值,科学准确性难度高于沟通能力,核心应关注交付、宣称与产物的联合分布。
中文摘要 AI 辅助
现有的科学人工智能基准大多是为了便于评分而设计的,比如多项选择题、带有参考解决方案的精心设计的智能体任务,或是具有已知生成结构的模拟器。而真实的科学请求则截然不同:它们的规格不完整、附带附件,且缺乏真实值。我们报告K-Bench 01,这是一个基于K-Dense Web上实时用户流量的首轮请求构建的评估基准,通过九个前沿模型在相同的沙箱环境中端到端运行,共产生1602次完成的智能体运行。三名盲态语言模型法官依据包含八个维度的评分标准对每次运行进行评分,其中评分标准的第8级锚定条件要求法官判断领域科学家是否会接受该工作并仅做少量编辑。结果显示,没有任何模型能在三名法官的评判下均达到该阈值。gpt-5.6-sol的合并均值最高,为8.04,但其95%置信区间[7.80, 8.23]跨越了阈值,且三名法官中有两名将claude-opus-5排在首位。因此,我们将系统的排序作为可复现的量,将绝对水平视为评估工具的属性,并将排名首位的情况视为未解决。在所有39934条评分判断中(每个评估的八个维度得分加上整体综合得分,排除不适用的条目),47.6%的得分低于8分的阈值。评分标准各维度的难度并不均匀:在九个模型的所有评估中,科学准确性的平均得分为6.22,而沟通能力的平均得分为7.33,二者分母相同且方向一致。最主要的失败标签是过度宣称,占所有评估的31.4%。我们认为,科学智能体的有效衡量指标并非排行榜排名,而是所交付内容、所宣称内容以及所生成产物的联合分布。
英文摘要
Benchmarks for scientific artificial intelligence are mostly written to be scored: multiple-choice questions, curated agent tasks with reference solutions, or simulators with a known generative structure. Real scientific requests arrive differently. They are underspecified, they carry attachments, and they lack ground truth. We report K-Bench 01, an evaluation built from first-turn requests sampled from live user traffic on K-Dense Web and run end to end by nine frontier models in identical sandboxes, yielding 1,602 completed agent runs. Three blinded language-model judges scored every run against an eight-dimension rubric. On a rubric whose 8-anchor is defined as work a domain scientist would accept with minor edits, no model clears the line under all three judges. gpt-5.6-sol has the highest pooled mean, 8.04, but its 95% interval [7.80, 8.23] spans the threshold, and two of the three judges rank claude-opus-5 first instead. We therefore report the ordering of systems as the reproducible quantity, the absolute level as an attribute of the instrument, and the top of the table as unresolved. Across all 39,934 scored judgments -- the eight dimension scores plus a holistic overall for each assessment, excluding not-applicable cells -- 47.6% fall below the 8-point threshold. Difficulty is not uniform across the rubric: scientific accuracy averages 6.22 against 7.33 for communication, on identical denominators and in the same direction within every one of the nine models. The single leading failure tag is overclaiming, on 31.4% of assessments. We argue that the informative quantity for scientific agents is not a leaderboard position but the joint distribution of what was delivered, what was claimed, and what artifacts were produced.