面向多任务密集预测的原型记忆任务-状态自适应
Task-State Adaptation with Prototype Memory for Multi-Task Dense Prediction
浏览论文内容
中文总结 AI 辅助
提出MemMTL多任务密集预测框架,结合原型记忆与稀疏路由,在NYUD-v2等数据集上评估其预测质量、计算成本及各模块贡献。
中文摘要 AI 辅助
视觉基础骨干网络为密集预测提供了强表征,但单一共享特征仍需满足不同的、依赖图像的任务自适应需求。我们提出MemMTL,这是一种多任务密集预测框架,可从全局视觉上下文估计紧凑的任务状态,并通过可学习的任务-状态原型记忆对其进行优化。优化后的状态被转换为任务条件专家logits,并在所有任务共享的局部专家库上进行稀疏top-k选择前,与token级logits结合。独立的任务无关残差库提供通用自适应路径,两条路径在任务特定预测前各添加一次到骨干特征中。我们在NYUD-v2和PASCAL-Context上指定匹配的评估协议,采用SAM 3和ViT-L骨干网络,以测量预测质量、计算成本以及任务-状态条件、原型检索和稀疏路由的贡献。本工作草案中的数值记录早于该规范实现,必须重新生成才能支持实证结论。
英文摘要
Vision foundation backbones provide strong representations for dense prediction, yet a single shared feature still needs to support tasks with different, image-dependent adaptation requirements. We propose MemMTL, a multi-task dense prediction framework that estimates a compact task state from global visual context and refines it through a learnable task-state prototype memory. The refined state is converted into task-conditioned expert logits and combined with token-level logits before sparse top-$k$ selection over a local expert bank shared by all tasks. A separate task-agnostic residual bank provides a common adaptation path, and both paths are added once to the backbone feature before task-specific prediction. We specify a matched evaluation protocol on NYUD-v2 and PASCAL-Context with SAM 3 and ViT-L backbones to measure predictive quality, computational cost, and the contributions of task-state conditioning, prototype retrieval, and sparse routing. The numerical record in the present working draft predates this canonical implementation and must be regenerated before it can support empirical claims.
发表机构
- Tsinghua University(清华大学)
- University of California, Merced(加州大学默塞德分校)
- Shanghai Jiao Tong University(上海交通大学)
机构由 AI 辅助整理,请以论文原文为准。