arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

空间动作审查:用于审计电子显微镜中语言到动作交接的可视化分析仪表盘

Spatial Action Review: A Visual Analytics Dashboard for Auditing Language-to-Action Hand-offs in Electron Microscopy

Samia Mohinta, Albert Cardona

arXiv 2609.24470首次发表:更新:

发表机构

University of Cambridge; MRC Laboratory of Molecular Biology(剑桥大学; MRC分子生物学实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对多模态大模型在科学图像分析中语言答案与空间动作不一致导致的静默失败,提出空间动作审查可视化仪表盘,通过账本、风险图和审计视图揭示弱耦合,支持人机交接决策。

AI 中文摘要

多模态大语言模型(MLLMs)正越来越多地被探索作为科学图像分析的接口,其中视觉问答(VQA)响应可能与指导下游阶段的空间输出配对。监督者阅读语言答案,而下游工作流(如分割或区域审查)则使用点集输出。我们将从检查答案到依赖其点动作的这一转变称为语言到动作交接。当答案正确而配对的动作遗漏了下游所需标注对象时,就会发生静默失败,因此基于答案的监督会清除动作不可靠的区域。我们引入了空间动作审查(Spatial Action Review),这是一个用于在电子显微镜(EM)线粒体分析中审计这种失败模式的可视化分析仪表盘。它通过答案-动作账本、任务-数据集风险图和图像区域审计视图将配对的答案-动作记录联系起来,将聚合模式与图像证据连接起来,同时可调节的动作可靠性门控支持重新审计。审查以人机交接结束,监督者记录动作是被接受、升级、在更严格的门控下保留,还是标记为模型修订。在来自EM适配的Qwen3-VL案例研究运行的541个图像区域中,点动作在54.4%的具有正确VQA响应的记录中未通过门控,27.4%的所有记录是静默失败。正确答案仅与可靠动作的概率提高5.8个百分点相关,自举区间跨越零;答案正确性与对象覆盖之间的点二列相关系数为0.061。这种弱耦合在753个匹配图像区域的五种模型条件下持续存在。空间动作审查使答案-动作不匹配可见,并将其与图像证据和记录的决策联系起来,然后MLLM输出进入自主科学工作流。

英文摘要

Multimodal large language models (MLLMs) are increasingly explored as interfaces for scientific image analysis, where a visual question-answering (VQA) response may be paired with a spatial output that guides a downstream stage. A supervisor reads the language answer, while a downstream workflow such as segmentation or region review consumes the point-set output. We call this transition from inspecting the answer to relying on its point action the language-to-action hand-off. A silent failure occurs when the answer is correct while the paired action misses annotated objects needed downstream, so answer-based oversight clears a region whose action is unreliable. We introduce Spatial Action Review, a visual analytics dashboard for auditing this failure mode in electron microscopy (EM) mitochondria analysis. It links paired answer-action records through an answer-action ledger, a task-by-dataset risk map, and an image-region audit view, connecting aggregate patterns to image evidence while an adjustable action-reliability gate supports re-audit. The review ends in a human-AI hand-off, where a supervisor records whether the action is accepted, escalated, held under a stricter gate, or flagged for model revision. Across 541 image regions from an EM-adapted Qwen3-VL case-study run, point actions fail the gate in 54.4% of records with a correct VQA response, and 27.4% of all records are silent failures. A correct answer is associated with only a 5.8-percentage-point higher probability of a reliable action, with a bootstrap interval spanning zero; the point-biserial correlation between answer correctness and object coverage is 0.061. This weak coupling persists across five model conditions on 753 matched image regions. Spatial Action Review makes answer-action mismatches visible and ties them to image evidence and a recorded decision before MLLM outputs enter autonomous scientific workflows.

CommentsAccepted at the IEEE VIS 2026 Workshop on Visual Analytics in the Age of Autonomous Science (VAxAutoSci)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑