门控激活引导:减少医学问答中的谄媚行为与幻觉
Gated Activation Steering for Reducing Sycophancy & Hallucination in Medical Question Answering
浏览论文内容
中文总结 AI 辅助
本研究提出门控激活引导方法,通过推理时干预结合幻觉与谄媚行为的针对性引导,在40亿参数模型上提升医学问答抗压力能力,表现接近千亿参数模型
中文摘要 AI 辅助
谄媚行为与幻觉是大型语言模型(LLMs)在各领域中持续存在的失效模式,而在临床问答场景中,这种问题尤为关键——其响应必须基于提供的上下文,且能抵御用户压力。幻觉会引入上下文不支持的信息,而谄媚行为则会使模型在受到用户质疑时放弃先前正确的答案。现有方法如基于提示的安全措施和持续激活引导,往往单独处理这些行为,或在多轮对话中广泛应用干预,可能不必要地损害原本正确的响应。为在单一框架内解决这些局限,本研究采用推理时干预(ITI),通过从对比临床对中学习幻觉与谄媚行为的独立引导方向,并将其应用于经因果验证的注意力头,以共同控制两种行为。运行时,行为特异性门控决定何时需要干预:幻觉组件缓解无依据的主张,谄媚组件缓解由用户压力导致的答案偏移。我们在基于电子健康记录(EHR)数据的临床问题上评估该框架,同时保持模型权重冻结。在所有评估设置中,我们进行了15900次模型响应运行。对于40亿参数模型的600条压力轨迹,未应用引导的模型在570个案例中屈服,而门控引导帮助其中551个案例维持更久。它在压力下坚守立场的表现可与拥有超过1000亿参数的模型相媲美,表明有针对性的推理时引导可提升鲁棒性,无需在每一轮都进行干预。
英文摘要
Sycophancy and hallucination are persistent failure modes of Large Language Models (LLMs) across domains. However, it becomes particularly consequential in clinical question answering, where responses must remain grounded in the provided context and robust to user pressure. Hallucination can introduce information that is unsupported by the context, while sycophancy can cause a model to abandon a previously correct answer when challenged by the user. Existing approaches, such as prompt-based safeguards and always-on activation steering, often address these behaviors separately or apply interventions broadly across turns, which can unnecessarily deteriorate responses that were already correct. To address these limitations within a single framework, we employ Inference Time Intervention (ITI) to jointly control both behaviors by learning separate steering directions for hallucination and sycophancy from contrastive clinical pairs and applying them to causally verified attention heads. During runtime, behavior-specific gates then determine when intervention is needed: the hallucination component mitigates unsupported claims, while the sycophancy component mitigates answer shifts caused by user pressure. We evaluate this framework on clinical questions grounded in EHR data while keeping the model weights frozen. Across all evaluation settings, we conducted 15,900 model-response runs. Across 600 pressure trajectories for the 4-billion-parameter model, the unsteered model caved in 570 cases. At the same time, gated steering helped it last longer in 551 of them. It held its ground under pressure at levels comparable to those of models with more than 100 billion parameters, showing that targeted inference-time steering can improve robustness without intervening at every turn.