arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.15045cs.CLcs.AI

魔镜魔镜告诉我:小型指令语言模型中的提示回显

Mirror, Mirror on the Wall: Prompt Echoing in Small Instruct Language Models

Inez Okulska, Bartosz Naskręcki, Jan Piotrowski, Tomasz Steifer

首次发表
浏览论文内容

中文总结 AI 辅助

本研究探究小型指令语言模型中的提示回显现象,发现其虽与训练数据部分重叠,但主要由模型的归纳头机制驱动,而非数据泄露所致。

中文摘要 AI 辅助

提示回显是指令语言模型的一种公认失败模式,即模型在未收到特定指令要求的情况下,不是生成响应,而是镜像复制所提供的提示。这一现象是模型泄露其训练数据集内容的迹象,还是由内部归纳/复制机制的失调行为所致?我们研究了来自不同系列(Gemma、Llama、Qwen、SmolLM 和 OLMo)的小型语言模型的提示回显现象,并表明回显提示很可能与训练数据集存在部分重叠,但该现象主要由模型的归纳头驱动。

英文摘要

Prompt echoing is a recognized failure mode of instruct language models, in which a model instead of generating a response, mirrors the provided prompt, even though it did not receive a specific instruction to do so. Is this phenomenon a sign of the model leaking the content of its training dataset, or is it rather caused by a misaligned behavior of the internal induction/copying mechanisms? We investigate prompt echoing small language models from different families (Gemma, Llama, Qwen, SmolLM and OLMo) and show that echoing prompts are likely to have partial overlap with the training dataset but the phenomenon is primarily driven by the model's induction heads.

发表机构

  • Centre for Credible AI, Warsaw University of Technology(华沙理工大学可信人工智能中心)
  • Adam Mickiewicz University Poznan(波兹南亚当·密茨凯维奇大学)
  • Institute of Fundamental Technological Research, Polish Academy of Sciences(波兰科学院基础技术研究所)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑