arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SyRuP:通过大语言模型解码中的奖励引导预测增强系统提示遵循

SyRuP: Enhancing System-Prompt Following via Reward-Guided Prediction in LLM Decoding

Seoyeon Kim, Minjae Kang, Jaehyung Kim

arXiv 2607.23991首次发表:更新:

发表机构

Yonsei University(延世大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究如何增强大语言模型对系统提示的遵循,提出SyRuP框架,通过训练交叉注意力奖励头产生令牌级遵循分数,推理时结合多种信号对候选重新排序,实验表明该方法能有效提升系统提示遵循度且开销适度。

AI 中文摘要

大语言模型(LLMs)越来越多地通过系统提示进行控制,这些提示规定了角色、风格、格式和安全要求。然而,模型仅通过上下文学习隐式遵循这些提示,对于复杂或组合式提示可能不够。现有方法通常需要模型调整或响应级重新排序,限制了其在轻量级推理时控制的实用性。我们引入了SyRuP,这是一个解码时框架,用于在保持基础大语言模型冻结的同时提高系统提示遵循度。SyRuP从系统提示条件偏好对中训练一个交叉注意力奖励头,将系统提示视为单独内存以产生令牌级遵循分数。在推理时,SyRuP通过将基础对数与学习到的奖励信号和捕获系统诱导对数偏移的可选对比信号相结合,对基础大语言模型的前k个候选进行重新排序。在系统提示遵循基准上的实验表明,SyRuP在适度推理开销下始终优于提示和解码时基线。这些结果表明,显式令牌级指导是可靠系统提示遵循的有效且实用机制。

英文摘要

Large Language Models (LLMs) are increasingly controlled through system prompts that specify roles, formats, and safety requirements. However, models follow these prompts only implicitly through in-context learning, which can be insufficient for complex or compositional prompts. Existing approaches often require model tuning or response-level reranking, limiting their practicality for lightweight inference-time control. We introduce SyRuP, a decoding-time framework for improving system-prompt adherence while keeping the base LM frozen. SyRuP trains a cross-attention reward head from system-prompt-conditioned preference pairs, treating the system prompt as a separate memory to produce token-level adherence scores. At inference, SyRuP reranks the base LM's top-k candidates by combining base logits with both the learned reward signal and a contrastive signal that captures system-induced logit shifts. Experiments on system-prompt following benchmarks show that SyRuP consistently outperforms prompting and decoding-time baselines with moderate inference overhead. These results suggest that explicit token-level guidance is an effective and practical mechanism for reliable system-prompt following.

CommentsEMNLP 26 Main, 27 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑