arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.21320stat.MLcs.LG

对角化注意力用于个体化回归:潜在行定位与预测

Diagonalized Attention for Individualized Regression: Latent-Row Localization and Prediction

Borui Peng, Liwei Lin, Feifei Wang, Long Feng

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出对角化注意力机制,用于矩阵值数据中样本特定行的定位与个体化回归,理论保证恢复潜在行,实验验证预测与定位性能。

中文摘要 AI 辅助

现代文本和图像表示通常是矩阵值形式的,其行对应于词元、图像块或其他局部特征向量。预测信息往往是稀疏但样本特定的,这使得具有共同支撑集的经典稀疏回归方法难以适应这种异质性。本文形式化了一个针对矩阵值协变量的个体化稀疏回归框架,在该框架中,每个观测值拥有自己感兴趣的行,而相关的回归效应在总体中是共享的。为了估计该模型,我们引入了一种对角化注意力机制,该机制使用查询-键分数来定位样本特定的信号行,并使用值矩阵进行下游回归。所提出的方法具有与样本量无关的参数维度,并且可以在没有响应的情况下识别新观测值的感兴趣行。我们建立了存在性定理,表明在适当的分数分离和集中条件下,单头和多头对角化注意力模型能够以高概率恢复潜在行,从而给出预测风险界。因此,我们的理论从统计学角度解释了基于注意力的评分如何在异质性矩阵值数据中定位样本特定信号。模拟实验表明,在不同样本量、维度和信号基数下,该方法在回归和错误设定的分类中均表现出强大的预测和定位能力。真实情感分析显示,分类准确率提高且词元选择具有可解释性。

英文摘要

Modern text and image representations are often matrix-valued, with rows corresponding to tokens, patches, or other local feature vectors. Predictive information is often sparse but sample-specific, making classical sparse regression methods with a common support poorly suited to this heterogeneity. This paper formalizes an individualized sparse regression framework for matrix-valued covariates in which each observation has its own rows of interest, while the associated regression effects are shared across the population. To estimate this model, we introduce a diagonalized attention mechanism that uses query--key scores to localize sample-specific signal rows and a value matrix for downstream regression. The proposed method has a parameter dimension independent of sample size and can identify rows of interest for new observations without their responses. We establish existence theorems showing that, under suitable score-separation and concentration conditions, single-head and multi-head diagonalized attention models recover the latent rows with high probability, yielding prediction risk bounds. Our theory therefore provides a statistical explanation of how attention-based scoring localizes sample-specific signals in heterogeneous matrix-valued data. Simulations demonstrate strong prediction and localization in regression and misspecified classification across varying sample sizes, dimensions, and signal cardinalities. Real sentiment analyses show improved classification accuracy and interpretable token selection.

发表机构

  • University of Hong Kong(香港大学)
  • Renmin University of China(中国人民大学)

机构由 AI 辅助整理,请以论文原文为准。

↑