Reliable Unlearning Harmful Information in LLMs with Metamorphosis Representation Projection
机构 * Peking University(北京大学) ; Tsinghua University(清华大学)
Comments 10 pages, 9 figures, Under review as a full paper at AAAI 2026. A preliminary version is under review at the NeurIPS 2025 Workshop on Reliable ML from Unreliable Data