arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向数据高效RL对齐的宪法网格工具(C-Guard)

C-Instrument: Automating RL Data Generation and Hillclimbing with a Constitution-Grid Instrument

Xianling Zhang

arXiv 2608.00180首次发表:更新:

AI 中文总结

针对RL对齐中目标冲突与数据效率问题,提出C-Guard宪法网格工具及C-LIM可学习性评分,优化了拒绝率与学习影响,相关代码与宪法已开源。

AI 中文摘要

强化学习(RL)对齐中普遍存在目标冲突问题,且在该场景下实现数据高效训练颇具挑战性。利用RL训练安全防护器意味着需要优化两个相互冲突的目标:既要识别真实危害,又不能拒绝良性提示。我们的研究发现,过度拒绝率从22.4%降至12.8%,而针对对抗性攻击的拒绝不足则会悄无声息地从0.27恶化至0.33。我们提出了C-Guard,这是一种用于生成RL训练数据的宪法网格工具,以及C-LIM,这是一种用于决定每个网格单元操作(剪枝、致密化、修正、扩展)的单元可学习性评分。C-LIM会在任何训练预算被消耗之前标记出无效数据区域:187个无针对性的行数据未带来任何增益,而我们的方法将同一区域的学习影响从0.733提升至0.80。相关代码与宪法已开源。

英文摘要

Conflicting objectives are general in RL alignment, and training on them data-efficiently is hard. Training a safety guard with RL means optimizing two objectives that conflict: catch real harm, and do not refuse benign prompts. Our finding is that over-refusal improves 22.4% to 12.8%, while under-refusal on adversarial attacks silently worsens 0.27 to 0.33. We present C-Instrument, a constitution-grid data instrument that generates the RL training data, and C-LIM, a per-cell learnability score that decides each cell's move: prune, densify, amend, expand. C-LIM flags the dead-weight data region before any training budget is spent: 187 untargeted rows had bought zero gain, and our method lifts the same region's learning impact 0.733 to 0.80. Code and the constitution are open-sourced.

CommentsPublished at COLM 2026 Efficient Reasoning

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑