arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

将大语言模型与图卷积网络集成用于半监督图像分类

Integrating Large Language Models and Graph Convolutional Networks for Semi-Supervised Image Classification

Camila Piscioneri Magalhães, Lucas Pascotti Valem

arXiv 2607.09104首次发表:更新:

发表机构

Institute of Mathematics and Computer Science (ICMC); University of São Paulo (USP)(数学与计算机科学研究所; 圣保罗大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对图像分类中图构建难题,该研究将大语言模型与图卷积网络集成,利用视觉语言模型生成文本描述,经大语言模型处理得到语义相似性分数,指导kNN图边修剪,提升了半监督图像分类准确率。

AI 中文摘要

随着图像数据的增加,标注数据集成本高且耗时,半监督方法如GCN应运而生。将GCN应用于图像分类的主要挑战之一是图构建。多数研究基于预训练深度学习骨干的特征向量相似性构建图。大语言模型在捕捉高级语义方面能力显著,但与GCN集成用于图像分类的研究较少。本文使用视觉语言模型生成文本图像描述,经大语言模型处理估计连接图像间的语义相似性分数,指导kNN和互反kNN图的边修剪,过滤掉语义不相关邻居。实验结果表明利用大语言模型进行图优化可提高分类准确率,特别是对于kNN图和某些骨干。

英文摘要

While the growing availability of image data has driven significant advances, labeling datasets remains costly and time-consuming. Therefore, semi-supervised approaches such as Graph Convolutional Networks (GCNs), which learn from both labeled and unlabeled data, have emerged as a promising solution. One of the primary challenges in applying GCNs to image classification is graph construction, since, unlike in citation networks or similar domains, images typically do not come with a predefined structural representation. For visual data, most studies construct graphs based on the similarity between feature vectors from pretrained deep learning backbones, typically by employing kNN or reciprocal kNN algorithms. Although Large Language Models (LLMs) have shown remarkable capability in capturing high-level semantics, their integration with GCNs for image classification remains underexplored. Aiming to fill this gap, our approach uses a Vision Language Model (VLM) to generate textual image descriptions, which are then processed by an LLM to estimate semantic similarity scores between connected images. These scores guide the pruning of edges in kNN and reciprocal kNN graphs, filtering out semantically irrelevant neighbors. Experimental results reveal that leveraging LLMs for graph refinement can improve classification accuracy, particularly for kNN graphs and some backbones. The source code is publicly available at http://gcnllm.lucasvalem.com.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑