PersoDPO: Scalable Preference Optimization for Instruction-Adherent, Persona-Grounded Dialogue via Multi-LLM Evaluation
PersoDPO: 通过多LLM评估实现可扩展的偏好优化
机构 * School of Computing, Macquarie University(计算机学院,麦考瑞大学)
专题命中 后训练与偏好优化 :preference optimization(title,abstract);LLM(title);large language model(abstract);language model(abstract)
AI总结 PersoDPO通过多LLM评估实现可扩展的偏好优化,提升对话系统的人格导向和上下文连贯性。
Comments Accepted at WISE 2025 Conference