arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

大型语言模型能否将粘贴的工件与用户语音区分开?未标记提示接缝处的吸收现象

Can LLMs Separate Pasted Artifacts from User Speech? Absorption at Unmarked Prompt Seams

Sugam Panthi, Muhaiminul Yeamin, Rabab Abdelfattah

arXiv 2610.04210首次发表:更新:

发表机构

AIMS Lab, The University of Southern Mississippi(南密西西比大学AIMS实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究LLMs在未标记提示接缝处吸收后续用户评论的现象,提出SEAM基准,发现裸换行吸收率高,边界标记可降低但无法消除。

AI 中文摘要

大型语言模型(LLMs)将每条用户消息视为纯文本,即使该消息组合了来自不同来源的文本。例如,用户可能将文本粘贴到提示中,并直接在下方继续输入评论。我们研究吸收现象:模型将后续的用户评论视为粘贴文本的一部分,并将其返回在编辑后的文本中。即使用户无意让评论成为该文本的一部分,这种情况也会发生。现有的指令-数据分离基准会告知模型哪些文本是指令,哪些是数据,然后测试其是否遵守该分离。这些基准并未测试在未标记的粘贴之后的无害用户语音。我们引入了SEAM,一个包含300个编辑示例的受控基准。每个示例在六种匹配条件下进行测试,这些条件改变了粘贴文本与后续用户语音之间边界的表达方式。在20个模型中,裸换行处的吸收率介于7.7%至66.7%之间。添加空行并未显著降低任何模型中的吸收率,而边界标记在20个模型中的19个中降低了吸收率。与粘贴文本匹配的评论(例如在代码后输入的代码注释)在20个模型中的17个中被显著更频繁地吸收。模型通常无法将粘贴的材料与后续的用户语音区分开,而显式边界可以减少但无法消除这种失败。

英文摘要

Large language models (LLMs) receive each user message as plain text, even when it combines text from different sources. For example, a user may paste text into a prompt and keep typing a comment directly below it. We study absorption: a phenomenon where the model treats a trailing user comment as part of the pasted text, returning it inside the edited text. This happens even though the user did not intend the comment to become part of that text. Existing instruction-data separation benchmarks tell the model which text is instruction and which is data, then test whether it obeys that separation. They do not test harmless user speech following an unmarked paste. We introduce SEAM, a controlled benchmark of 300 editing examples. Each example is tested under six matched conditions that vary how the boundary between pasted text and later user speech is expressed. Across 20 models, absorption at a bare newline ranges from 7.7% to 66.7%. Adding a blank line does not significantly reduce absorption in any model, while boundary markers reduce it in 19 of 20 models. Comments that fit the pasted text, such as a code comment typed after code, are absorbed significantly more often in 17 of 20 models. Models often fail to separate pasted material from later user speech, and explicit boundaries reduce but do not remove this failure.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑