arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.18056cs.CV

位置锚点调优:面向预训练点云Transformer的高效适配

Position Anchor Tuning: Towards Efficient Adaptation of Pre-Trained Point Cloud Transformers

Zheng Liu, Xin Gao, Jinchao Zhu, Gao Huang

首次发表
浏览论文内容

中文总结 AI 辅助

针对预训练点云Transformer适配中推理效率被忽视的问题,提出位置锚点调优(PAT),通过令牌聚合-扩展对和基础共享低秩适配,在保持性能的同时显著降低计算开销和可训练参数。

中文摘要 AI 辅助

参数高效微调(PEFT)近来已成为将预训练点云Transformer适配到各种下游任务的关键研究方向。尽管现有方法在实现高参数效率的同时取得了优异的微调性能,但它们忽略了推理效率。为解决这一问题,本文提出了一种新颖的PEFT方法,称为位置锚点调优(PAT)。由于多头注意力(MHA)和前馈网络(FFN)是预训练Transformer中计算密集的模块,PAT通过令牌聚合-扩展对来降低其计算成本。每一对包含一个令牌聚合模块(TAM)和一个令牌扩展模块(TEM)。对于MHA和FFN模块,TAMs基于3D空间中的位置锚点从其输入令牌中提取代表性令牌。这些提取的令牌(而非原始输入令牌)被模块处理,从而减少了参与计算的令牌数量。随后,TEMs将学习到的表示传播回原始输入令牌。由于TAMs仅负责捕获任务特定的表示,进一步引入了基础共享低秩适配(BSLoRA),使其仅用少量可训练参数就能有效学习此类表示。在广泛使用的基准上的大量实验表明,PAT在显著降低计算开销和可训练参数的同时,性能与最先进方法相当。

英文摘要

Parameter-efficient fine-tuning (PEFT) has recently emerged as a pivotal research direction for adapting pre-trained point cloud transformers to diverse downstream tasks. Although existing methods achieve excellent fine-tuning performance with high parameter efficiency, they ignore inference efficiency. To tackle this problem, a novel PEFT method termed position anchor tuning (PAT) is proposed in this paper. As multi-head attention (MHA) and feed-forward network (FFN) are computation-heavy blocks in pre-trained transformers, PAT decreases their computational cost through token aggregation-expansion pairs. Each pair comprises a token aggregation module (TAM) and a token expansion module (TEM). For MHA and FFN blocks, TAMs extract representative tokens from their input tokens based on position anchors in 3D space. These extracted tokens, rather than the original input tokens, are processed by the blocks, thereby reducing the number of tokens involved in computation. Then, TEMs propagate the learned representations back to the original input tokens. Since TAMs are solely responsible for capturing task-specific representations, base-sharing low-rank adaptation (BSLoRA) is further introduced to enable them to learn such representations effectively with only a small number of trainable parameters. Extensive experiments on widely used benchmarks demonstrate that PAT performs comparably to state-of-the-art methods while incurring significantly lower computational overhead and fewer trainable parameters.

发表机构

  • University of Science and Technology Beijing(北京科技大学)
  • Beijing Engineering Research Center of Industrial Spectrum Imaging(北京工业光谱成像工程技术研究中心)
  • Nankai University(南开大学)
  • Tsinghua University(清华大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑