PROPS: Progressively Private Self-alignment of Large Language Models
PROPS: 大型语言模型的逐步隐私自对齐
AI总结 PROPS通过多阶段隐私保护对齐框架,在保护偏好标签隐私的同时提升LLM对齐效果,实现更高的胜率。
Comments Accepted in the Transactions on Machine Learning Research (TMLR), 2025
Journal ref Transactions on ML Research (TMLR) 2025