arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.21315cs.CL

提示-模型交互达到不动点:一种确定性、无任务的结构读出——以及失败的因子分解

Prompt-Model Interaction Reaches the Fixed Points: A deterministic, task-free structural readout -- and the factorizations of it that failed

Nicolás Vera Zúñiga

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对提示-模型交互的无任务读出,发现9个token的条件可改变不动点结构类别与模型排序,提出的现象学因素无法解释该效应,解释单元为提示-模型对。

中文摘要 AI 辅助

已有研究证实,提示的效果并非提示本身的属性:针对某一模型优化的提示在另一模型上性能会下降,且在中性重新格式化下排序会发生变化。该证据与任务准确率相关,但无法说明这种交互是关于任务机制的事实还是关于条件分布本身的事实。我们在一种不含任务的读出方式上展开研究:从96个起始点统计短窗口argmax映射的不动点结构,该映射为x_{t+1}=argmax_x p(x | x_{t-1}, x_t)。该映射是确定性的,不存在任何可被帮助或损害的因素,且仅存在于短窗口中——6个模型中有4个在窗口16时完全失去该结构——因此此处的所有内容都与模型如何读取片段有关。我们得到两个结果:第一,交互以全幅度达到该读出方式:9个token的条件将不动点分数在其大部分范围内移动,改变了四类结构类别,并重新排序模型,而具有60.5个IFEval分数的指令微调对该类别的改变为零;第二,我们提出的所有因素都无法解释该现象。前缀长度不成立:该效应非单调。四个现象学因素——散文与标记、通用方向、双向性、指令抗性——每个因素在提出后一次运行内就被撤回,因样本扩大而消失。而最接近的机制解释(早期token的注意力汇聚主导)仅能预测5个模型中2个的移位符号(属于随机概率),同时长度-内容交叉显示该效应在真实文本上成立,在我们探测的均匀随机输入上不成立,因此我们处于该机制的适用范围之外,而非与其相悖。一个固定的9-token前缀驱动4个模型趋向0,2个模型趋向1;双向性在分布内起始中保留。在该读出方式上,解释的单元是提示-模型对。我们发现的反复出现的错误有一个名称:将具有某种形状的标准应用于无变化空间的量。

英文摘要

That a prompt's effect is not a property of the prompt is established: prompts optimised for one model degrade on another, and rankings reorder under neutral reformatting. That evidence is about task accuracy, which cannot say whether the interaction is a fact about task machinery or about the conditional distribution itself. We ask on a readout with no task in it: the fixed-point structure of the short-window argmax map x_{t+1} = argmax_x p(x | x_{t-1}, x_t), censused from 96 starts. It is deterministic, so nothing can be helped or hurt, and it exists only at short windows -- four of six models lose it entirely by window 16 -- so everything here concerns how a model reads a fragment. Two results. First, the interaction reaches this readout at full magnitude: nine tokens of conditioning move the fixed-point fraction across most of its range, change a four-way structural class, and reorder models, while instruction tuning worth 60.5 IFEval points moves the class by zero. Second, nothing we proposed carries it. Prefix length fails: the effect is not monotone. Four phenomenological factors -- prose-versus-markup, a universal direction, bidirectionality, instruct-resistance -- were each withdrawn within one run of being proposed, dissolved by widening the sample. And the nearest mechanistic account, attention-sink dominance of early tokens, predicts the sign of the shift on 2 of 5 models -- chance -- while a length-by-content cross shows it holds on real text and fails on our probe's uniformly random input, so we are outside its regime, not against it. One fixed nine-token prefix drives four models toward 0 and two toward 1; the bidirectionality survives in-distribution starts. On this readout the unit of explanation is the prompt-model pair. The recurring error it caught in us has a name: a criterion with a shape applied to a quantity with no room to vary.

补充信息

↑