Understanding Impact of Human Feedback via Influence Functions
机构 * KAIST(韩国科学技术院) ; Columbia University(哥伦比亚大学)
专题命中 后训练与偏好优化 :RLHF(abstract,comments);LLM(abstract);large language model(abstract);language model(abstract)
Comments Accepted at ACL 2025, Source code: https://github.com/mintaywon/IF_RLHF
Journal ref Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics 63 (2025) 27471-27500