When Do Language Models Endorse Limitations on Human Rights Principles?
语言模型何时会支持人权原则的限制?
机构 * Max Planck Institute for Intelligent Systems(马克斯·普朗克智能系统研究所) ; Jinesis AI Lab, University of Toronto & Vector Institute(Jinesis AI实验室,多伦多大学及向量研究所) ; University of Michigan(密歇根大学) ; University of Copenhagen(哥本哈根大学) ; EuroSafeAI
AI总结 本文研究了语言模型在人权原则限制上的偏见,发现模型在不同语言和提示下表现出系统性差异,揭示了AI在高风险互动中的伦理挑战。
Comments EACL Findings 2026