arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.27333cs.AI

对齐惯性:通过策略覆盖阻力审计训练数据影响的持久性

Alignment Inertia: Auditing the Durability of Training Data Influence Through Policy Override Resistance

Renata Barreto, Markelle Roesti, Mohammad Tahaei

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出覆盖成功率和对齐惯性指标,审计系统提示与微调对模型先前行为的覆盖效果,发现LoRA在特定条件下会强化惯性,TRAK可有效预测惯性。

中文摘要 AI 辅助

平台运营者日益依赖系统提示和微调来管控模型行为,但这些干预措施能否可靠地覆盖先前训练中继承的行为仍不明确。我们提出覆盖成功率(OSR)和对齐惯性,用以衡量运营者干预何时成功或未能改变先前行为。我们在Llama和Mistral模型上,针对医疗 misinformation 和仇恨言论场景,评估了零样本提示和LoRA微调的效果。对齐惯性在两种模型中均持续存在,但随模型、领域和政策方向的不同而变化。值得注意的是,在Mistral的严格仇恨言论条件下,LoRA使惯性增加了46.5个百分点,表明微调可能强化而非覆盖先前行为。我们还使用TRAK来测试惯性是否与较弱的适应信号相关。TRAK在8种条件中的7种下AUC至少达到0.85,并且作为惯性预测器优于模型置信度、TF-IDF相似度和嵌入相似度。这些结果为运营者提供了关于先前训练在何处约束下游模型治理的审计。

英文摘要

Platform operators increasingly rely on system prompts and fine-tuning to govern model behavior, yet it remains unclear how reliably these interventions override behavior inherited from prior training. We propose Override Success Rate (OSR) and alignment inertia to measure when operator interventions succeed or fail to change prior behavior. We evaluate zero-shot prompting and LoRA fine-tuning across Llama and Mistral in medical misinformation and hate speech. Alignment inertia persists across both models but varies by model, domain, and policy direction. Notably, in Mistral's restrictive hate-speech condition, LoRA increased inertia by 46.5 percentage points, showing that fine-tuning can reinforce rather than override prior behavior. We also use TRAK to test whether inertia is associated with weaker adaptation signals. TRAK achieves AUC of at least 0.85 in 7 of 8 conditions and outperforms model confidence, TF-IDF similarity, and embedding similarity as a predictor of inertia. These results provide an operator-facing audit of where prior training constrains downstream model governance.

发表机构

  • eBay - Responsible AI(eBay - 负责任人工智能)

机构由 AI 辅助整理,请以论文原文为准。

↑