USPLIT-VQA:用于视觉问答的U形分割学习与贡献感知加权聚合
USPLIT-VQA: U-Shaped Split Learning for Visual Question Answering with Contribution-Aware Weighted Aggregation
- Southern Illinois University Carbondale(南伊利诺伊大学卡本代尔分校)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
提出USPLIT-VQA,一种U形分割学习框架,结合贡献感知加权聚合(CAWA)实现隐私保护VQA,在多个数据集上降低内存和通信开销,并有效抑制恶意客户端影响。
AI中文摘要:
视觉问答(VQA)系统联合解释图像和自然语言查询,在众多领域具有显著应用前景,但用户数据的隐私敏感性构成了根本性障碍。集中式训练需要访问所有数据,而联邦学习要求每个客户端托管完整模型。我们提出USPLIT-VQA,一种用于隐私保护VQA的U形分割学习框架,其中每个客户端保留初始层和分类头,而服务器托管计算量大的中间层,将原始输入和标签保留在客户端设备上。我们进一步引入贡献感知加权聚合(CAWA),一种基于梯度相似性的客户端评分机制,旨在减少恶意更新的影响。在四个VQA数据集(VQA-RAD、SLAKE、PathVQA和VizWiz)上使用两种骨干网络的实验表明,在评估的固定分割下,Custom模型相比联邦学习获得了准确率提升,而BiomedCLIP的准确率有所下降,同时客户端内存减少高达5.8倍,通信量减少高达10.8倍。在存在一个恶意客户端的情况下,CAWA将攻击者的影响降低了超过98%,而在更高腐败水平下的实验揭示了其局限性。重建实验进一步表明,在所评估的攻击下,反演质量较低。
英文摘要:
Visual Question Answering (VQA) systems, jointly interpreting images and natural language queries, hold significant promise across many domains, yet the privacy-sensitive nature of user data creates a fundamental barrier. Centralized training requires access to all data, while federated learning requires each client to host the full model. We propose USPLIT-VQA, a U-shaped split learning framework for privacy-preserving VQA in which each client retains the initial layers and the classification head while the server hosts the computationally heavy intermediate layers, keeping raw inputs and labels on the client device. We further introduce Contribution-Aware Weighted Aggregation (CAWA), a gradientsimilarity-based client scoring mechanism designed to reduce the influence of malicious updates. Experiments on four VQA datasets (VQA-RAD, SLAKE, PathVQA, and VizWiz) with two backbones show accuracy gains over Federated Learning for the Custom model and reduced accuracy for BiomedCLIP under the evaluated fixed split, alongside client memory reductions of up to 5.8X and communication reductions of up to 10.8X. With one malicious client, CAWA reduces the attacker's influence by over 98%, while experiments at higher corruption levels identify its limitations. Reconstruction experiments further show lower inversion quality under the evaluated attacks.