Hedging and Non-Affirmation: Quantifying LLM Alignment on Questions of Human Rights
对冲与非肯定:量化大语言模型在人权问题上的对齐
机构 * Google Deepmind(谷歌DeepMind) ; Massachusetts Institute of Technology(麻省理工学院) ; Independent Researcher(独立研究员) ; Google(谷歌) ; AI Accountability Lab, Trinity College Dublin(都柏林圣三一学院人工智能问责实验室)
专题命中 其他LLM :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.AI
AI总结 研究通过系统框架量化LLM在不同群体身份上的对冲与非肯定行为,发现群体身份是主要影响因素,通过引导和正交化技术可有效缓解偏差。