arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于注意力的多任务计算表示

Attention-based representations for multi-task computation

Daniel Hsu, Mingyue Xu

arXiv 2608.04243首次发表:更新:

发表机构

Columbia University; Purdue University(哥伦比亚大学; 普渡大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对两类多任务场景,建立了多头注意力层所需头数量的下界并构造出匹配下界的结构,相关结果可推广到任意对称布尔函数。

AI 中文摘要

多头注意力层生成的向量表示可支持多个下游任务。我们针对两个简单且具体的多任务场景,建立了所需注意力头数量的界。第一个场景中,需要一种向量表示,使得线性预测器能够计算给定列表中的最小数和最大数;已知在此场景下,两个具有小嵌入维度和比特精度级别的注意力头就足够,我们证明单个注意力头需要指数级更高的嵌入维度或精度级别。第二个场景中,需要一种向量表示,使得多项式阈值函数能够计算给定n位字符串的异或(XOR);对于n=2的情况,该场景与第一个场景类似,因为异或可通过对编码两位的与(AND)和或(OR)的向量表示使用线性函数轻松计算。我们观察到,n位异或所需的注意力头数量与多项式次数的乘积至少为n,并且我们构造了达到该下界的多头注意力层。这些结果可推广到任意(对称)布尔函数,其中界由阈值次数给出。

英文摘要

Multi-head attention layers produce vector representations that support multiple downstream tasks. We establish bounds on the number of heads required in two simple and concrete multi-task scenarios. In the first scenario, a vector representation is sought so that linear predictors can compute both the smallest and largest numbers in a given list. In this case, it is known two attention heads with small embedding dimension and bit precision level suffice. We prove that a single attention head requires exponentially higher embedding dimension or precision level. In the second scenario, a vector representation is sought so that a polynomial threshold function can compute the XOR of a given string of $n$ bits. This scenario is analogous to the first one for $n=2$, since XOR is readily computed by a linear function using a vector representation that encodes both the AND and the OR of the two bits. We observe that $n$-bit XOR requires the product of the number of heads and the polynomial degree to be at least $n$, and we construct multi-head attention layers that match this lower bound. These results generalize to arbitrary (symmetric) Boolean functions, where the bound is given in terms of the threshold degree.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑