Operationalizing Pluralistic Values in Large Language Model Alignment Reveals Trade-offs in Safety, Inclusivity, and Model Behavior
在大型语言模型对齐中操作多元价值观揭示了安全、包容性和模型行为之间的权衡
机构 * Technical University of Munich(慕尼黑技术大学) ; Stanford University(斯坦福大学) ; Cornell University(康奈尔大学)
专题命中 后训练与偏好优化 :large language model(title,abstract);language model(title,abstract);LLM(abstract);preference optimization(abstract)
AI总结 本研究探讨了在大型语言模型对齐中融入多元价值观的影响,发现不同群体偏好和设计选择对模型行为和安全性有显著影响。