arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

BIABench:评估AI代理在真实生物图像分析任务中的表现

BIABench: Evaluating AI agents on real-world bioimage analysis tasks

Zixuan Pan, Davide Panzeri, Lukas Johanns, Marilin Moor, Yu Zhou, Hedi Peterson, Yiyu Shi, Jianxu Chen

arXiv 2609.34274首次发表:更新:

AI 中文总结

BIABench是一个由16项真实生物学任务组成的基准,用于评估AI代理端到端生物图像分析能力,发现复杂三维或时序任务表现不佳且结果不稳定。

AI 中文摘要

人工智能(AI)代理有望实现生物图像分析的自动化,但目前尚无基准测试能够评估它们是否能够端到端地完成真实世界的分析任务。这类分析对代理而言颇具挑战,因为二维图像、三维体数据和时序序列往往过大,无法直接作为上下文读取,因此代理必须通过代码、专业软件和渲染视图来选择并执行分析。已发表的研究使这种能力变得可测试,因为每项研究都将原始图像与同行评审的结果配对。我们推出了BIABench,这是一个从已发表的生物学研究中重建的16项任务基准,这些任务保留了其科学问题、成像数据和真实标签。这些任务涵盖了从H&E组织学到单分子定位显微镜的11种分析子任务和模态。每次提交都会获得一个结果分数,该分数使用领域标准指标将输出文件与真实标签进行比较,以及一个过程分数,其中视觉语言模型根据专家编写的评分标准评判方法选择和质量控制。我们评估了多种语言模型上的通用型和生物学专用型代理,并对每项任务进行了重复运行。常规的二维任务解决得很好,但在某些增加第三维或时间轴的任务上,没有任何代理得分超过0.19。无论是生物学专业化、更强的模型还是详细的专家指令,都无法弥合这一差距。代理也不可靠,同一代理重复运行之间的分数差异大于不同代理之间的差异,而且在没有真实标签的情况下,无法通过过程分数或花费的时间区分正确运行和错误运行。BIABench及其数据和代码已公开开放,为评估并最终训练可靠的长周期生物图像分析代理提供了一个可验证的框架。

英文摘要

Artificial-intelligence (AI) agents hold promise for automating bioimage analysis, yet no benchmark evaluates whether they can carry out real-world analyses end to end. Such analyses are hard for agents because 2D images, 3D volumes and time-lapse sequences are often too large to read as context, so an agent must choose and run an analysis through code, specialized software and rendered views. Published studies make this capability testable, because each pairs raw images with a peer-reviewed result. We introduce BIABench, a benchmark of 16 tasks reconstructed from published biological studies that retain their scientific questions, imaging data and ground truth. The tasks span eleven analysis subtasks and modalities from H&E histology to single-molecule localization microscopy. Each submission receives an outcome score, which compares the output files with the ground truth using field-standard metrics, and a process score, in which a vision-language model judges method choice and quality control against an expert-written rubric. We evaluated general-purpose and biology-specific agents across several language models, with repeated runs of every task. Routine two-dimensional tasks were solved well, but on some tasks that added a third dimension or a time axis no agent scored above 0.19. Neither biological specialization, stronger models nor detailed expert instructions closed this gap. The agents were also unreliable, with scores varying more between repeated runs of one agent than between different agents, and without ground truth a correct run could not be told from a wrong one by its process score or by the time spent. Released openly with its data and code, BIABench provides a verifiable framework for evaluating, and eventually training, agents for reliable long-horizon bioimage analysis.

Comments41 pages, 6 figures, 11 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑