arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.00361cs.CVcs.AI

用于光学显微镜下颗粒与纤维表征的人工智能方法

Artificial Intelligence for the Characterization of Particles and Fibers by Optical Microscopy

Simiao Sun, Kenneth Ng, Lynn Lee, Astrid Harth, Asami Odate, Aggelos Katsaggelos, Manuel Ballester Matito, Nicholas Eastaugh, Marc Walton

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出一种带语义锚点的AI蒸馏框架,训练仅视觉输入的学生ViT模型,在颗粒与纤维显微镜表征任务中获80%伪类验证准确率等结果,实现更丰富可解释的表征,适用于相关分析场景。

中文摘要 AI 辅助

颗粒与纤维分散体系的光学显微镜观测涉及对受样本形态、化学成分、放大倍率及光照条件影响的细微视觉线索的解读。本文提出一种人工智能(AI)蒸馏框架,利用语义锚点从显微镜图像中提取语义丰富的图像嵌入。一个多模态教师模型将每张图像的视觉嵌入与代表光照模态、放大倍率及样本身份与形态的三个文本嵌入相结合,该文本嵌入由LongCLIP的扩展上下文文本编码器生成,最终得到一个2304维的块结构教师向量,其组成块在训练和推理过程中始终保持物理可解释性。我们训练了一个带有多层感知机(MLP)解码器的学生视觉Transformer(ViT),使其仅从图像中重构该教师向量,最小化平均绝对误差(L1)损失以确保与教师块结构的坐标级保真度。来自教师嵌入空间的HDBSCAN聚类生成的伪类上的交叉熵项作为防坍塌正则化器,在无需对比负样本挖掘的情况下强制类间分离。推理时,学生仅需图像输入,即可生成能恢复教师向量全部语义内容的紧凑嵌入。该框架在留一法最近邻检索下,伪类验证准确率约达80%,细粒度样本描述标签的Recall@1达75%。这些结果表明,语义锚定使仅视觉输入的学生模型能获得比仅图像训练更丰富、更可解释的表征,可直接应用于异质颗粒与纤维分散体系的检索、分类及探索性分析。

英文摘要

Optical microscopy of particle and fiber dispersions involves interpreting subtle visual cues influenced by specimen morphology, chemical composition, magnification, and illumination conditions. We introduce an artificial intelligence (AI) distillation framework that extracts semantically rich image embeddings from microscopy images using semantic anchors. A multimodal teacher combines each image's visual embedding with three text embeddings representing illumination modality, magnification, and specimen identity and morphology. Generated by LongCLIP's extended-context text encoder, this yields a 2304-dimensional block-structured teacher vector whose component blocks remain physically interpretable throughout training and inference. A student vision transformer (ViT) with a multi-layer perceptron (MLP) decoder is trained to reconstruct this teacher vector from the image alone, minimizing a mean absolute error (L1) loss that enforces coordinate-level fidelity to the teacher's block structure. A cross-entropy term over pseudo-classes derived from HDBSCAN clustering of the teacher embedding space acts as a collapse-prevention regularizer, enforcing inter-cluster separation without requiring contrastive negative mining. At inference, the student operates on image input alone, producing compact embeddings that recover the full semantic content of the teacher vector. The framework achieves approximately 80% pseudo-class validation accuracy and 75% Recall@1 on fine-grained specimen description labels under leave-one-out nearest-neighbor retrieval. These results demonstrate that semantic anchoring enables a vision-only student to acquire richer and more interpretable representations than image-only training, with direct applicability to retrieval, classification, and exploratory analysis of heterogeneous particle and fiber dispersions.

↑