arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.35490cs.CV

AHMAD:用于关键点检测的自适应混合多任务视觉学习与辅助蒸馏

AHMAD: Adaptive Hybrid Multi-task Vision Learning with Assisted Distillation for Keypoint Detection

Mohammad Mahdi, Nedyalko Prisadnikov, Yuqian Fu, Carmelo Scribano, Danda Pani Paudel, Luc Van Gool

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出AHMAD框架,统一五种视觉任务,通过共享编码器-解码器和轻量级投影器实现多任务学习,并引入知识蒸馏优化关键点检测效率,达到SOTA性能。

中文摘要 AI 辅助

通用多任务视觉模型旨在将多个视觉任务统一到一个框架内,从而实现更高效、更通用的学习。然而,处理涵盖密集和稀疏预测的多种视觉任务仍然具有挑战性,因为它们的输出结构本质上是不同的。在本文中,我们提出了AHMAD,一个简单而有效的通用多任务学习框架,它整合了不同的关键视觉任务:语义分割、实例分割、深度估计、关键点检测和对象检测。我们的方法将这五个任务纳入一个统一的结构:一个共享的编码器-解码器,带有几个轻量级的任务特定投影器。在多任务学习范式下,我们观察到了互补的性能提升,在COCO-val全景分割和语义分割上分别达到了最先进的PQ 53.1和mIoU 66.5。此外,对于自上而下的关键点检测,由于多次前向传播通常导致高计算开销,我们引入了一种基于知识蒸馏的方法,使得对整个图像只需一次前向传播,大大提高了效率。最终,我们的模型提供了一个轻量级但有效的通用多任务学习框架,在五个视觉任务上展示了强大的性能。

英文摘要

Generalist multitasking vision models aim to unify multiple vision tasks within a single framework, enabling more efficient and versatile learning. However, handling diverse vision tasks -- spanning dense and sparse predictions -- remains challenging due to their inherently varying output structures. In this paper, we propose AHMAD, a simple yet effective framework for generalist multitask learning that integrates different key vision tasks: semantic segmentation, instance segmentation, depth estimation, keypoint detection, and object detection. Our approach incorporates these five tasks into a unified structure: a shared encoder-decoder with several lightweight task-specific projectors. Under the multitask learning paradigm, we observed a complementary performance gain, achieving a state-of-the-art PQ of 53.1 and an mIoU of 66.5 for COCO-val panoptic and semantic segmentation, respectively. Additionally, for top-down keypoint detection, which typically incurs high computational overhead due to multiple forward passes, we introduce a knowledge distillation-based method that enables a single forward pass over the entire image, greatly improving efficiency. Ultimately, our model delivers a lightweight yet effective generalist multitask learning framework, demonstrating strong performance across five vision tasks.

↑