FedCoT-VQA:面向视频问答中思维链规划器的联邦学习与遗忘框架
FedCoT-VQA: A Federated Learning and Unlearning Framework for Chain-of-Thought Planners in VideoQA
浏览论文内容
中文总结 AI 辅助
提出FedCoT-VQA框架,通过规划器侧分区、服务器侧聚合和残差遗忘模块,在联邦学习中高效训练思维链规划器并支持客户端删除,提升接地质量4.45%,遗忘后反事实差距仅7.38%。
中文摘要 AI 辅助
思维链(Chain-of-Thought, CoT)规划器已成为视频问答(VideoQA)的一种有效设计,其中轻量级规划器首先生成中间推理步骤,以在答案预测之前引导时间证据选择。这种模块化特性使得基于CoT的VideoQA对联邦学习具有吸引力,因为只有规划器侧需要协作适应,而庞大的视觉-语言骨干网络可以保持固定。然而,在去中心化环境中,规划器不仅需要在异构客户端上高效训练,还需要支持后续的客户端删除请求。这具有挑战性,因为被删除客户端的影响既体现在模型参数中,也体现在规划器的推理轨迹行为上。我们提出了FedCoT-VQA,一个面向VideoQA中CoT规划器的联邦学习与遗忘框架。FedCoT-VQA由三个模块组成:规划器侧分区(PSP),它暴露了一个紧凑的共享残差适应空间,用于高效的联邦训练;服务器侧聚合(SSA),它聚合规划器侧更新,同时维护一个支持删除的贡献日志;以及残差遗忘模块(RUM),它通过保留客户端重放和选择性残差校正来近似仅保留的反事实规划器,无需完全重新训练。我们从联邦训练效用、联邦遗忘效用、遗忘质量和效率方面评估了FedCoT-VQA。结果表明,与当前的联邦方法相比,FedCoT-VQA保持了强大的联邦训练效用,将接地质量提高了高达4.45%。遗忘后,它保持了高准确率,并实现了仅7.38%的反事实差距。
英文摘要
Chain-of-Thought (CoT) planners have emerged as an effective design for VideoQA, where a lightweight planner first generates intermediate reasoning steps to guide temporal evidence selection before answer prediction. This modularity makes CoT-based VideoQA attractive for federated learning, since only the planner side needs collaborative adaptation while the heavy vision-language backbone can remain fixed. However, in decentralized settings, the planner must not only be trained efficiently across heterogeneous clients but also support later client deletion requests. This is challenging because deleted-client influence is reflected both in model parameters and the planner's reasoning-trace behavior. We present FedCoT-VQA, a federated learning and unlearning framework for CoT planners in VideoQA. FedCoT-VQA consists of three modules: planner-side partitioning (PSP), which exposes a compact shared-residual adaptation space for efficient federated training; server-side aggregation (SSA), which aggregates planner-side updates while maintaining a deletion-ready contribution log; and a residual unlearning module (RUM), which approximates the retained-only counterfactual planner through retained-client replay and selective residual correction, without full retraining. We evaluate FedCoT-VQA in terms of federated training utility, federated unlearning utility, forgetting quality, and efficiency. Results show that compared to current federated approaches, FedCoT-VQA preserves strong federated training utility, improving grounding quality by up to 4.45%. After unlearning, it retains high accuracy and achieves a counterfactual gap of only 7.38%.
发表机构
- The Hong Kong Polytechnic University(香港理工大学)
- Hong Kong University of Science and Technology(香港科技大学)
- Pengcheng Laboratory(鹏城实验室)
机构由 AI 辅助整理,请以论文原文为准。