Sleight of Word基准:语言模型能否察觉自身输出被篡改?
Sleight of Word Benchmark: Can Language Models Notice If Their Own Output Was Tampered With?
浏览论文内容
中文总结 AI 辅助
该研究构建了Sleight of Word基准,测试19种开放权重语言模型对自身生成过程中单个词被替换的外部扰动的感知,从惊讶度指标和文本反应两方面评估模型的察觉能力。
中文摘要 AI 辅助
语言模型的输出可能在其生成过程中被外部篡改。因此可通过评估模型对这种外部扰动的感知来构建简单测试。基于此,本文构建了一个简单基准,在生成过程中持续用另一个词替换单个词,称该方法为Sleight of Word。研究从两个不同维度进行测量:与模型惊讶度相关的指标,以及对19种不同开放权重语言模型的文本反应评估。
英文摘要
The output of a Language Model can be tampered with \emph{while} the model is writing it. A simple test can thus be constructed by evaluating the model's perception of this external perturbation. In this spirit, a simple benchmark is built in which a single word is consistently substituted with another in the generation process. We call this method \emph{Sleight of Word}. Two distinct axes are measured: metrics that relate to the model's surprise, as well as an evaluation of the textual reaction for 19 different open-weight language models.
发表机构
- Cleo
机构由 AI 辅助整理,请以论文原文为准。