arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.01306cs.AIcs.CL

DAYJOB:面向长周期专业工作的基准测试

DAYJOB: A Benchmark for Long-Horizon Professional Work

Stephanie Finley, Liudas Panavas, Thomas Mikkelson, Cam Hinton, Stacey Ganss, Bradley Monton, Emily Kendall, Michelle Spradlin, Lydia Bye, Michael O'Brien, Laur… 展开作者

Stephanie Finley, Liudas Panavas, Thomas Mikkelson, Cam Hinton, Stacey Ganss, Bradley Monton, Emily Kendall, Michelle Spradlin, Lydia Bye, Michael O'Brien, Lauren Ylvisaker, Derek Ray, Suhaas Garre, Sushant Mehta, Edwin Chen

首次发表
浏览论文内容

中文总结 AI 辅助

DAYJOB是一个由专业人士构建的130个长周期任务基准,评估AI智能体在医疗和金融领域的专业工作能力,最强模型通过率不足25%,揭示了智能体在前提验证和错误输入处理上的缺陷。

中文摘要 AI 辅助

专业工作通常始于一个简短的请求,专业人士需要自行确定所需内容、哪些文件重要,以及请求的前提是否成立。我们推出了DAYJOB,一个由医疗保健(50个)和金融(80个)领域专业人士构建的包含130个任务的基准测试。这些任务预计平均耗时在医疗保健领域为13.6小时,在金融领域为16.6小时。每个任务都是一个容器化的Harbor环境,配有专家制定的二元标准评分细则(每任务中位数分别为47.5和57.5),由智能体评审员对交付文件进行评判,只有满足所有标准才算通过。在来自13个开发者的30种模型配置中,最强的Claude Opus 5.5在医疗保健任务中通过率为24.7%,在金融任务中为23.9%,而中位数配置的通过率分别为0.6%和2.5%。在案例研究中,智能体会接受与记录相矛盾的前提,并将错误的输入带入本应一致的分析中。我们发布了所有医疗保健任务、80个金融任务中的50个、评估工具包以及排行榜。

英文摘要

Professional work often starts with a brief request that leaves the professional to work out what is needed, which documents matter, and whether the request's premise holds. We introduce DAYJOB, a benchmark of 130 tasks built by professionals in healthcare (50) and finance (80). The tasks are estimated to take a professional 13.6 hours on average in healthcare and 16.6 in finance. Each task is a containerized Harbor environment with an expert rubric of binary criteria (median 47.5 and 57.5 per task) that an agentic judge applies to the delivered files, and an attempt passes only if it meets every criterion. Across 30 model configurations from 13 developers, the strongest, Claude Opus 5.5, passes 24.7% of healthcare and 23.9% of finance attempts, and the median configuration passes 0.6% and 2.5%. In case studies, agents accept premises that the record contradicts and carry wrong inputs through otherwise consistent analyses. We release all healthcare tasks, 50 of the 80 finance tasks, the evaluation harness, and the leaderboard.

发表机构

  • Surge AI

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑