arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于鲁棒跨数据集图像分类的MLLM路由异质集成模型

MLLM-Routed Heterogeneous Ensembles for Robust Cross-Dataset Image Classification

Daniel Perkins, John Squires, Janou Milligan, Chandra Raskoti, Linda Ungerboeck

arXiv 2608.13463首次发表:更新:

发表机构

The Bredesen Center for Interdisciplinary Research and Graduate Education; University of Tennessee Knoxville(布雷德森跨学科研究与研究生教育中心; 田纳西大学诺克斯维尔分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对跨数据集图像分类泛化难题,提出ARMDIL模型,通过MLLM智能体动态路由图像到适配的视觉骨干,兼具竞争力、高适应性与可解释性,为通用视觉系统奠定基础。

AI 中文摘要

现代图像分类模型在针对单一任务特定数据集进行训练时表现出色,但往往难以跨不同领域和难度级别进行泛化。我们提出了ARMDIL,即一种结合大型语言模型(LLM)的多领域图像分类自适应路由模型。ARMDIL是一种集成模型,它使用多模态大型语言模型(MLLM)智能体将每张图像动态路由到最合适的视觉骨干网络。我们的多样化集成模型采用了卷积神经网络(ResNets)、自监督表征学习器(SSL)以及视觉语言模型(VLMs),每个模型均基于由多个具有不同分布和特征的图像数据集构建的统一标签空间进行训练。实证评估揭示了各架构在不同视觉领域的独特能力与弱点。至关重要的是,我们表明ARMDIL能有效平衡这些权衡,其表现可与基于专门训练的路由模型相媲美。此外,它通过简单的提示修改即可整合新信息,大幅提升了适应性,同时通过自然语言推理轨迹增强了可解释性。这些跨数据集图像分类领域的进展为构建更可靠的通用视觉系统(如AI助手和自主机器人)铺平了道路。

英文摘要

Modern image classification models excel when trained on single task-specific datasets but often struggle to generalize across domains and difficulty levels. We propose ARMDIL, an Adaptive Router for Multi-Domain Image Classification with LLMs. ARMDIL is an ensemble that uses a multimodal large language model (MLLM) agent to dynamically route each image to the most suitable vision backbone. Our diverse ensemble employs convolutional neural networks (ResNets), self-supervised representation learners (SSL), and vision language models (VLMs), each trained on a unified label space constructed from multiple image datasets with differing distributions and characteristics. Empirical evaluations illuminate the distinct capabilities and vulnerabilities of each architecture across disparate visual domains. Crucially, we show that ARMDIL effectively navigates these tradeoffs, performing competitively with specialized training-based routers. Furthermore, it drastically improves adaptability by allowing new information to be integrated via simple prompt modifications, while enhancing interpretability through natural language reasoning traces. These advances in cross-dataset image classification pave the way for more reliable general-purpose vision systems such as AI assistants and autonomous robots.

Comments15 pages (single column), 4 figures, 7 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑