arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

FedSSMCoOp:基于SSM编码器的轻量级联邦提示学习用于少样本分类

FedSSMCoOp: SSM Encoders for light-weight Federated Prompt Learning for Few-shot Classification

Ankita Das, Ambarish Parthasarathy, Sumohana S. Channappayya, C. Krishna Mohan

arXiv 2610.09907首次发表:更新:

发表机构

Indian Institute of Technology Hyderabad(印度海得拉巴理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对联邦生物医学少样本分类中余弦相似度无法捕获跨模态交互的问题,提出FedSSMCoOp框架,利用SSM编码器仅优化提示更新,实现轻量级多模态学习,性能稳定且平均轻量化1.96倍。

AI 中文摘要

视觉-语言模型(VLMs)凭借其各自领域中包含的互补信息,在广泛的下游视觉任务中展现出强大的性能。尽管性能有所提升,但这些方法大多依赖于使用余弦相似度度量来对齐这些领域,这无法在分类阶段之前捕获令牌级结构和跨模态交互。这在联邦约束下的生物医学应用中尤为关键,因为在这些应用中,数据共享受到限制,每个站点的标记数据稀缺,且各机构之间的数据差异很大,导致显著的统计异质性。为了克服这一问题,我们提出了FedSSMCoOp,一个联邦少样本图像分类框架,能够在保护数据隐私的同时实现多模态学习。借助基于SSM的Vision Mamba和Cross Mamba块,并且通过在联邦设置中仅优化软提示和通信提示的更新,该框架优先考虑了计算和性能。重要的是,这消除了使用外部大型语言模型(LLM)进行特征对齐的需要。该框架进一步在各种生物医学图像数据集上进行训练和评估,并对其性能进行了评估。所提出的框架相对于基线提供了稳定的性能,并且平均轻量化1.96倍。相应的脚本将很快发布。

英文摘要

Vision-Language Models (VLMs) have shown strong performance across a wide range of downstream vision tasks, thanks to the complementary information contained in the respective domains. Despite the performance gains, most of these approaches rely on aligning these domains using the cosine similarity metric, which fails to capture token-level structure and cross-modal interactions prior to the classification stage. This is especially critical in biomedical applications under federated constraints, where data sharing is restricted, labeled data is scarce at each site, and it differs widely across institutions, leading to substantial statistical heterogeneity. To overcome this issue, we propose FedSSMCoOp, a federated few-shot image classification framework that enables multimodal learning while preserving data privacy. With the help of the SSM-based Vision Mamba and Cross Mamba blocks, and by optimizing only the soft-prompt and communication-prompt updates in the federated setting, the framework prioritizes both computation and performance. Importantly, this eliminates the need to use an external Large Language Model (LLM) for feature alignment. The framework is further trained and evaluated on various biomedical image datasets, and its performance is assessed. The proposed framework delivers stable performance relative to the baselines and is, on average, 1.96 times lighter. The corresponding script will be made available soon.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑