D-STEER - Preference Alignment Techniques Learn to Behave, not to Believe -- Beneath the Surface, DPO as Steering Vector Perturbation in Activation Space
D-STEER - 通过偏好对齐技术学习行为而非信念 -- 在表象之下,DPO作为激活空间中的低秩引导机制
机构 * IIIT Delhi(印度理工学院德里分校) ; Microsoft(微软) ; Apple (USA)(苹果(美国)) ; Google (USA)(谷歌(美国)) ; Pragya Lab, BITS Pilani, K. K. Birla Goa Campus(普拉吉亚实验室,比斯科大学,K.K. 布尔拉果阿校区)
专题命中 偏好对齐 :alignment(title,abstract);DPO(title,abstract);分类 cs.LG
AI总结 D-STEER研究揭示DPO通过引导激活空间中的低秩机制实现行为对齐,而非改变模型内部信念。