InSPO: Unlocking Intrinsic Self-Reflection for LLM Preference Optimization
InSPO:解锁LLM偏好优化的内在自反思
机构 * George Washington University(乔治·华盛顿大学)
专题命中 后训练与偏好优化 :preference optimization(title,abstract);LLM(title);large language model(abstract);language model(abstract)
AI总结 InSPO通过解锁LLM的内在自反思能力,提升偏好优化的鲁棒性和人类对齐性。