arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

适配开放权重多模态大语言模型以生成电子显微镜分割的点提示

Adapting Open-Weight MLLMs to Generate Point Prompts for Electron Microscopy Segmentation

Samia Mohinta, Albert Cardona

arXiv 2609.14080首次发表:更新:

发表机构

University of Cambridge; MRC Laboratory of Molecular Biology(剑桥大学; MRC分子生物学实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究首次验证开放权重多模态大语言模型可依据自然语言生成点提示以驱动电子显微镜分割,经微调后性能显著提升,并展现出跨数据集迁移能力。

AI 中文摘要

诸如microSAM之类的可提示模型能够根据点提示对电子显微镜(EM)图像进行分割,但自动化需要在不依赖用户输入的情况下生成提示。我们探究开放权重的多模态大语言模型(MLLMs)是否能够通过返回坐标给冻结的分割器,从自然语言请求中生成这些提示。为此,我们将三个线粒体数据集中的掩膜转换为训练示例,将图像和指令与质心坐标配对,然后在冻结MLLM骨干和microSAM的同时训练LoRA适配器。我们发现,经过监督微调和奖励优化后,Qwen3-VL的分割AP$_{50}$达到0.736,而未适配时仅为0.247,同时自动提示生成(APG)达到0.773。此外,另外两个MLLM也有所提升,达到或超过了APG。与一个针对该线粒体任务达到AP$_{50}$ 0.904的监督质心热图检测器相比,Qwen3-VL更接近标注的点集和实例数量。而且,在两个公共数据集上训练可以迁移到未见过的第三个数据集,而在全部三个数据集上训练则可以迁移到独立的EM体积。鲁棒性测试显示,在未见过的自然语言请求表述下性能稳定,同时坐标可以被第二个分割器重用。据我们所知,这是首个关于开放权重MLLMs作为EM点生成器的可行性研究,为定位和掩膜解码之间提供了一个可检查的、语言引导的链接。

英文摘要

Promptable models such as microSAM segment electron microscopy (EM) images from point prompts, but automation requires generating prompts without user input. We ask whether open-weight multimodal large language models (MLLMs) can generate them from natural-language requests by returning coordinates to a frozen segmenter. To that end, we convert masks from three mitochondria datasets into training examples, pairing images and instructions with centroid coordinates, then train LoRA adapters while freezing the MLLM backbone and microSAM. We find that Qwen3-VL reaches segmentation AP$_{50}$ $0.736$ after supervised fine-tuning and reward optimization, up from $0.247$ without adaptation, while automatic prompt generation (APG) achieves $0.773$. In addition, two other MLLMs improve, reaching or exceeding APG. When compared with a supervised centroid-heatmap detector that reaches AP$_{50}$ $0.904$ for this mitochondria task, Qwen3-VL more closely matches the annotated point set and instance counts. Moreover, training on two public datasets transfers to an unseen third, while training on all three transfers to an independent EM volume. Robustness tests show stable performance under unseen formulations of the natural-language request, while the coordinates can be reused by a second segmenter. To our knowledge, this is the first feasibility study of open-weight MLLMs as EM point generators, providing an inspectable, language-directed link between localization and mask decoding.

CommentsAccepted at the BioImage Computing (BIC) Workshop at ECCV 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑