arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

关于支持保持对齐和有界过滤的局限性

On the Limits of Support-Preserving Alignment and Bounded Filtering

Aryan Dutt, Rui Mao, Anupam Chattopadhyay

arXiv 2607.18295首次发表:更新:

AI 中文总结

研究在大语言模型中,重塑基础模型输出分布的对齐方案与有界安全过滤器结合能否消除有害行为。通过形式化设置并分析,发现有界过滤可能无法消除所有有害输出,实证评估显示有害输出率始终高于零。

AI 中文摘要

我们研究重塑基础模型输出分布的对齐方案与有界安全过滤器相结合,能否使现代大语言模型中有害行为的概率降至零。近期研究表明基于偏好的对齐下有害行为可能持续,外部过滤在最坏情况下计算困难。我们在黑盒、白盒和统计查询访问下,用支持保持对齐算子和有界过滤算法形式化此设置,分析其逼近理想消除器的能力。基于此框架,我们提供计算和信息论论据表明在这些约束下,有界过滤可能无法消除基础模型分布支持的所有有害输出。为实证评估这些限制,我们在有界黑盒、白盒和统计查询过滤器下,对一系列通过OpenRouter访问的先进开放权重和托管大语言模型进行分析,使用来自精心策划的网络安全场景和PKU - SafeRLHF的对抗性提示。跨模型、过滤器类别和查询预算,估计的有害输出率随额外过滤计算而降低,但始终稳定在零以上,表明存在持续的经验性危害下限。

英文摘要

We study whether alignment schemes that reshape a base model's output distribution, combined with bounded safety filters, can drive the probability of harmful behavior to zero in modern large language models. Recent research suggests that harmful behaviors can persist under preference-based alignment and that external filtering can be computationally hard in the worst case, but it remains unclear whether practical alignment pipelines that largely preserve internal representations can eliminate harmful behavior entirely rather than merely suppressing its most visible forms. We formalize this setting using support-preserving alignment operators together with bounded filtering algorithms under black-box, white-box, and statistical-query access, and analyze their ability to approximate an ideal eliminator that removes all harmful mass. Building on this framework, we provide computational and information-theoretic arguments indicating that, under these constraints, bounded filtering may fail to eliminate all harmful outputs supported by the base model's distribution. To evaluate these limits empirically, we analyze a range of state-of-the-art open-weight and hosted LLMs accessed via OpenRouter under bounded black-box, white-box, and statistical-query filters on adversarial prompts drawn from curated cybersecurity scenarios and PKU-SafeRLHF. Across models, filter classes, and query budgets, the estimated harmful-output rate decreases with additional filtering compute but consistently plateaus above zero, suggesting a persistent empirical harm floor.

CommentsWithdrawn to allow a complete rewrite and substantial revision of the theoretical analysis and experiments

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑