大规模提示设计:格式、指令数量和上下文长度如何影响大语言模型中的指令遵循和幻觉
Prompt Design at Scale: How Format, Instruction Count, and Context Length Shape Instruction Adherence and Hallucination in Large Language Models
浏览论文内容
中文总结 AI 辅助
研究大语言模型提示设计中格式、指令数量和上下文长度对指令遵循及幻觉的影响,通过在合成语料库上的两个控制实验评估五个模型,揭示了不同因素下的表现及规律,还发布了相关工具和结果。
中文摘要 AI 辅助
从业者在进行提示设计时,对于如何格式化指令和上下文(如markdown、纯文本、散文或表格)、系统提示在合规性降低前能携带多少同步指令以及模型在召回率和诚实度降低前能容纳多少上下文,几乎没有可控证据。我们在一个无污染的合成语料库(“Veyra之书”,8780个唯一命名实体,可从固定种子确定性再生)上进行了两个控制实验,涉及这三个因素,并对五个模型进行评估。实验1测量随着规则数量N从10增长到160时指令遵循的衰减,实验2测量在2k到512k令牌上下文阶梯上相同四种格式下的召回准确率、虚假前提谄媚和无事实编造情况。结果表明,随着N增加,完美响应率下降,不同模型和格式下有不同表现,上下文长度增加时召回率先保持高位后下降,编造从未发生,谄媚可忽略不计,但拒绝回答率大幅上升,且预注册格式顺序不成立,令牌开销也影响格式偏好。我们发布了完整工具、语料库生成器和原始结果(VeyraBench)。
英文摘要
Practitioners make three prompt-design decisions with almost no controlled evidence behind them: how to format instructions and context (markdown, plain text, prose, or tabular), how many simultaneous instructions a system prompt can carry before compliance degrades, and how much context a model can hold before recall and honesty degrade. We report two controlled experiments crossing all three factors on one held, contamination-free synthetic corpus (the "Book of Veyra," 8,780 uniquely-named entities, deterministically regenerable from a fixed seed), evaluated across five models. Experiment 1 (960 calls/model) measures instruction-following decay as rule count N grows from 10 to 160, crossed with four formats and system-prompt vs. user-turn placement. Perfect-response rate collapses to zero by N=80 for every model, format, and placement. Placement produces effects at least as large as format at N=160 in most models, but the direction is model-specific. No model shows a reliable markdown advantage; one 35B model favors plain text instead. Experiment 2 (5,520 calls/model) measures recall accuracy, false-premise sycophancy, and absent-fact fabrication across a 2k-to-512k-token context ladder in the same four formats. Recall stays near ceiling through 64-128k tokens, then degrades sharply and format-dependently: one model's accuracy spread reaches 48 points at 128k tokens. Fabrication never occurs (0/5,760 probes), and sycophancy stays negligible (<=8.3%). What rises sharply near each model's context ceiling is outright refusal to answer (0% to 79-90%), distinct from sycophancy or fabrication. Neither pre-registered format ordering holds, and token overhead (+22% to +37% over plain text) further changes which format is preferable where accuracy spread is genuine. We release the full harness, corpus generator, and raw results (VeyraBench): https://github.com/iNetanel/veyrabench
发表机构
- Machine Human Intelligence Lab (MHIL)(机器人类智能实验室)
机构由 AI 辅助整理,请以论文原文为准。