Improving LLM Safety and Helpfulness using SFT and DPO: A Study on OPT-350M
机构 * Saarland University(萨尔兰大学)
专题命中 后训练与偏好优化 :SFT(title,abstract);LLM(title,comments);language model(abstract);RLHF(abstract)
Comments 17 pages, 3 figures. Code and dataset available at https://github.com/PiyushWithPant/Improving-LLM-Safety-and-Helpfulness-using-SFT-and-DPO