AI 中文总结
研究针对大语言模型推理中现有引导向量方法的问题,提出概率概念感知引导框架,通过概念驱动检索和概率校准,在保留模型能力的同时实现可控、安全的语义偏差引导。
AI 中文摘要
引导向量(SVs)是大语言模型(LLMs)推理时的一种干预技术,通过在推理过程中向中间激活添加特定概念的方向向量来引导生成过程。然而,现有SV方法常产生破坏可解释性和细粒度控制的表示不一致行为。本文提出概率概念感知引导(PCS)框架用于LLM推理,通过概念驱动的引导向量检索和概率强度校准,在保留原始任务能力的同时提供可控、面向安全的语义偏差。
英文摘要
Steering vectors (SVs), an inference-time intervention technique for large language models (LLMs), guide the generation process by adding a concept-specific direction vector to intermediate activations during inference. However, existing SV methods frequently yield representation-incoherent behaviors that undermine interpretability and fine-grained control, largely because prior work has focused on binary positive-negative steering evaluation while employing discrete clustering metrics that fail to capture the continuous spectrum of semantic alignment. In this work, we present the Probabilistic Concept-Aware Steering (PCS) framework for LLM inference. PCS preserves original task competence while providing controllable, safety-oriented semantic bias through concept-driven steering-vector retrieval and probabilistic strength calibration.