arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.14174cs.LGcs.CLcs.SD

继承的注意力头:音频语言模型通过其文本骨干的注意力追踪说话者,且一种注意力质量排序会检索出不同的集合

Inherited Heads: Audio language models track speakers with their text backbone's attention, and an attention-mass ranking retrieves a different set

  • Cornell University(康奈尔大学)

机构由 AI 辅助整理,请以论文原文为准。

Bojro Das

AI总结:

本研究揭示音频语言模型可通过文本骨干的注意力头追踪说话者,添加偏置可高精度重定向描述,且这些头多继承自文本模型,并比较了两种注意力头排序方法的效果。

AI中文摘要:

当被要求描述录音中六位说话者之一所谈论的内容时,音频语言模型在6%至16%的试验中描述了正确的说话者,低于随机猜测给出的16.7%。通过向一百个注意力头的注意力对数(logits)添加固定偏置,且无需训练,这些头仅占模型的一小部分(不到十分之一),即可在90.7%至99.0%的试验中将描述重定向到我们选择的任意说话者。这些头在很大程度上并非音频所特有。对构建音频模型所基于的纯文本语言模型,或同一系列已发布的模型,在任务的书面版本上进行排序,取其前一百个注意力头,并原样迁移:它们在80.8%至95.0%的试验中重定向了音频模型,而选择过程中完全不涉及音频。音频和文本注意力头集合在100个头中共享66至74个,而随机情况下大约共享20个,且仅共享部分就几乎重现了全部转向效果。但这并不能证明共享是这些头起作用的原因:从同一发现的一百个头中抽取相同大小的子集,其表现几乎一样好,而我们的三个模型均无法区分这两种解释。第二个发现涉及如何找到这些头。根据注意力头在被询问片段上放置的注意力量(如既有评分所做)或根据其注意力随问题移动的量(如每头归一化变体所做)对头进行排序,得到的排名前一百的头在我们的三个模型间分别共享69、37和4个头。在Ultravox中,它们共享4个头,既有评分的头在69.7%的试验中使输出无法被评判者归入任何片段,而无干预时为40.0%,归一化变体为1.0%。这是六个分支中的一个;在其他五个分支上,既有评分的转向效果优于随机抽取。

英文摘要:

Asked to describe what one of six speakers in a recording talks about, audio language models describe the right one on 6 to 16% of trials, below the 16.7% a guess would give. Adding a fixed bias to the attention logits of a hundred heads, under a tenth of the model's and with no training, redirects the description to whichever speaker we choose, on 90.7% to 99.0% of trials. Those heads are largely not specific to audio. Rank the text-only language model an audio model was built from, or a released model of the same family, on a written version of the task, take its top hundred heads, and carry them over unchanged: they redirect the audio model on 80.8% to 95.0% of trials, with nothing about audio entering the selection. The audio and text head sets share 66 to 74 of 100 where chance would give about 20, and the shared part alone reproduces almost all of the steering. What that does not show is that sharing is what makes the heads work: an equal-sized draw from the same discovered hundred does nearly as well, and none of our three models separates the two explanations. A second finding concerns how such heads are found. Ranking heads by how much attention they place on the segment asked about, as an established score does, or by how much of their attention moves with the question, as a per-head normalised variant does, gives top hundreds that share 69, 37 and 4 heads across our three models. In Ultravox, where they share 4, the established score's heads leave output the judge cannot place on any segment on 69.7% of trials, against 40.0% with no intervention and 1.0% for the normalised variant. That is one arm of six; on the other five the established score steers above a random draw.

补充信息

↑