使用状态空间模型进行长上下文演示选择
Long-Context Demonstration Selection Using State Space Models
- Northeastern University(东北大学)
- Emory University(埃默里大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出基于状态空间模型的长上下文演示选择方法,通过蒸馏Transformer和映射嵌入,实现线性推理成本,减少14.2倍FLOPs并提升6.48%准确率。
AI中文摘要:
我们研究演示选择问题,即从一组示例中选取子集,将其前置到语言模型的查询中。该问题与上下文学习和语言模型推理密切相关。由于Transformer模型的推理成本随序列长度呈二次方增长,在长上下文场景下,选择问题变得尤为棘手。本文通过构建状态空间模型(SSMs)来解决该问题,该模型仅需线性推理时间。我们的方法包含两种算法。第一种通过蒸馏(已训练的)Transformer模型来学习一组小型SSMs。我们将所有层划分为连续的组,然后为每组估计一个独立的状态空间模型,以复制相邻层内的输入-输出行为。第二种将蒸馏后的模型输出映射到一小部分token,并将这些嵌入用于下游应用中的演示选择。我们在合成和真实数据集上进行了大量实验以验证我们的方法。我们证明,蒸馏后的SSMs相对于真实输出的近似误差小于0.7%。在下游评估中,我们表明在多个文本分类和推理任务上,与基线演示选择方法相比,我们的方法将FLOPs减少了14.2倍,并将准确率提高了6.48%。
英文摘要:
We study the problem of demonstration selection, which involves selecting a subset of examples for prepending to a query to a language model. This problem is closely related to in-context learning and language model inference. Since the inference cost of a transformer model scales quadratically with sequence length, the selection problem becomes especially challenging in a long-context scenario. In this paper, we tackle this problem by building on state space models (SSMs), which require only linear inference time given the input. Our approach involves two algorithms. The first learns a small set of SSMs through distillation of a (trained) transformer model. We partition all the layers into consecutive groups. Then for each group, we estimate a separate state space model to replicate the input-output behavior within the adjacent layers. Second, we map the distilled model outputs to a small set of tokens, and apply these embeddings for demonstration selection in downstream applications. We perform extensive experiments in both synthetic and real-world datasets to validate our approach. We demonstrate that the distilled SSMs only incur an approximation error of less than $0.7\%$ relative to the true output. In downstream evaluation, we show that on several text classification and reasoning tasks, our approach reduces FLOPs by $14.2\times$ and improves accuracy by $6.48\%$ relative to baseline demonstration selection methods.