arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

识别导致幻影迁移的数据集偏差

Towards Identifying the Dataset Biases Causing Phantom Transfer

Jonas Jürß, Pietro Liò

arXiv 2609.14449首次发表:更新:

发表机构

University of Cambridge(剑桥大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出基于Sentence BERT嵌入的签名方法,用于识别导致幻影迁移的数据集偏差主题,在教师模型已知和未知时分别达到0.83和0.46的马修斯相关系数。

AI 中文摘要

近期研究表明,教师模型可以通过数据集将偏差迁移给学生模型,而该数据集中所有对该偏差的显式引用均已被过滤,且即使知道要寻找何种偏差,也没有数据层面的防御措施能够可靠地移除或检测该偏差。为了揭示这些偏差的隐藏痕迹,我们展示了一种基于Sentence BERT嵌入的简单签名,能够在攻击者使用的教师模型已知时,以0.83的马修斯相关系数识别此类偏差的主题;若教师模型未知,则相关系数为0.46。此外,我们观察到不同的教师模型似乎通过不同的词汇表达相同的偏差。

英文摘要

Recent work has shown that a teacher model can transfer a bias to a student through a dataset from which every explicit reference to that bias has been filtered out, and that none of the tested data-level defenses reliably removes or detects such a bias, even when the defender knows what to look for. To shed light on the hidden traces these biases leave, we embed a dataset's completions with Sentence-BERT, subtract the embeddings of clean reference completions, and compare the result to an open vocabulary of candidate topics. This simple signature identifies the topic of the bias with a Matthews correlation coefficient of 0.83 when the attacker's teacher model is known, and 0.46 when it is not. We also observe that different teacher models appear to express the same bias through different vocabulary.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑