arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

超越token级引导:通过跨系列表示转向实现专用大语言模型的推理时对齐

Beyond Token-Level Guidance: Inference-Time Alignment of Specialized LLMs via Cross-Family Representation Steering

Jin Gan, Xin Li, Jun Luo

arXiv 2608.30319首次发表:更新:

发表机构

College of Computing and Data Science, Nanyang Technological University; College of Cryptology and Cyber Science, Nankai University(南洋理工大学计算与数据科学学院; 南开大学密码与网络科学学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对专用LLM推理时对齐的缺陷,提出CREST方法,通过跨系列表示转向提升安全性,同时保留领域能力,在安全基准上比基线提升达22.2%。

AI 中文摘要

针对专用领域微调的大语言模型(LLMs)是高影响力应用的关键。推理时对齐可改善因专用微调而下降的安全性,且无需大量计算资源,作为一种易用的即插即用解决方案,补充了基于微调的方法。然而,现有推理时方法无法在不破坏领域能力的情况下可靠提升安全性。我们将根本原因识别为互补专长正交性:专用基础模型与通用领域引导模型的能力正交,导致引导信号对专用生成不可靠,主要表现为停止token干扰,即引导模型的延续倾向会覆盖基础模型的停止决策,将正确答案埋没在引导诱导的延续中。为解决该问题,我们提出CREST,一种推理时对齐方法,它利用从任意系列引导模型提取的安全方向来引导基础模型的隐藏表示,完全避开token级结构限制。CREST在专用微调削弱安全性的场景下提升了安全性,同时保留了领域特定能力和已良好对齐模型的安全性,在安全基准上比基线方法性能提升高达22.2%。我们的代码可在:this https URL获取。

英文摘要

Large language models (LLMs) finetuned for specialized domains represent crucial high-impact applications. Inference-time alignment improves safety degraded from specialization finetuning without requiring substantial computational resources, complementing finetuning-based methods with an easy-to-use, plug-and-play solution. However, existing inference-time methods fail to reliably improve safety without disrupting domain capability. We identify the root cause as complementary expertise orthogonality: specialized base models and general-domain guidance models have orthogonal competencies, making the guidance signal unreliable for specialized generation. This primarily manifests as stop token interference, where the guidance model's tendency toward continuation overrides the base model's decision to stop, burying correct answers under guidance-induced continuation. To address this problem, we propose CREST, an inference-time alignment method that steers base model hidden representations using safety directions extracted from a guidance model of any family, avoiding token-level structural limitations entirely. CREST improves safety where specialization has weakened it while preserving both domain-specific capability and the safety of already well-aligned models, outperforming baselines by up to 22.2\% on safety benchmarks. Our code is available at: https://github.com/DecayingSeart/CREST.

CommentsAccepted by EMNLP 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑