Corrigibility as a Singular Target: A Vision for Inherently Reliable Foundation Models
机构 * Independent Researcher(独立研究者) ; Machine Intelligence Research Institute(机器智能研究院)
专题命中 偏好对齐 :alignment(abstract);safety(abstract);分类 cs.AI、cs.CY、cs.LG
Comments Preprint. This work has been submitted to the Reliable and Responsible Foundation Models Workshop at ICML 2025 for review