arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.36492cs.CVcs.AIcs.LG

在连接组学中评估视觉语言模型的突触检测与校对能力

Benchmarking Vision-Language Models on Synapse Detection and Proofreading in Connectomics

  • Harvard University(哈佛大学)
  • Boston College(波士顿学院)
  • Nanyang Technological University(南洋理工大学)

机构由 AI 辅助整理,请以论文原文为准。

Yicong Li, Junjie Wang, Leander Lauenburg, Ella Hugie, Alexandra Irger, Wanhua Li, Donglai Wei, Hanspeter Pfister

AI总结:

本研究基准测试了视觉语言模型在连接组学中的突触检测与校对任务,发现零样本表现接近随机,而LoRA微调可使开源模型媲美专家模型,并在跨物种合并错误识别上超越专家模型。

AI中文摘要:

我们对视觉语言模型(VLMs)在连接组学中检查电子显微镜图像时注释者所做出的决策进行了基准测试:突触检测(存在性和极性)以及校对(分裂错误和合并错误)。对于突触检测,我们在零样本、四样本上下文学习和LoRA设置下,针对不同架构和规模的19个开源模型和2个闭源模型进行了评估,并与专家模型进行了比较,所使用的数据集由我们利用公共资源构建。对于校对,我们在ConnectomeBench2数据集上评估了3个开源模型和2个闭源模型,并进行了从果蝇和小鼠到人类和斑马鱼的跨物种迁移。大多数模型在零样本情况下表现接近随机水平;少量示例主要帮助了闭源模型和最大的开源模型。在数千个标签上进行LoRA微调使开源模型达到了与专家模型相当的水平。在评估未见过的物种时,经过最佳适配的VLM在识别合并错误方面优于在同一数据上训练的专家模型。该项目将在接收后公开发布。

英文摘要:

We benchmarked vision-language models (VLMs) on the decisions annotators take when inspecting electron microscopy images in connectomics: synapse detection (presence and polarity) and proofreading (split errors and merge errors). For synapse detection, we evaluated 19 open and 2 closed models across various architectures and sizes under zero-shot, four-shot in-context learning and LoRA settings, against specialist models, on datasets constructed by us using public resources. For proofreading, we evaluated 3 open and 2 closed models on the ConnectomeBench2 dataset, with cross-species transfer from fly and mouse to human and zebrafish. Most models were at chance zero-shot; a few examples helped mainly the closed and largest open ones. LoRA on a few thousand labels brought open models level with specialist models. When evaluated on unseen species, the best adapted VLMs outperformed specialist models trained on the same data in identifying merge errors. The project will be publicly available upon acceptance.

↑