The Algorithmic Unconscious: Structural Mechanisms and Implicit Biases in Large Language Models
算法无意识:大型语言模型中的结构机制与隐性偏见
专题命中 安全评测 :alignment(abstract);safety(abstract);分类 cs.CL、cs.CY
AI总结 本文探讨了大型语言模型中结构性机制与隐性偏见,揭示了分词、注意力等机制对偏见的影响,并提出技术诊所框架以促进AI基础设施的批判性利用。
Comments 18 pages, 5 figures, Extended version of a paper presented at the international conference 'Artificial Intelligence and Transformations of Information' (LOGOS/FLSH, Hassan II University of Casablanca, Morocco, December 2025), accepted for publication in LOGOS after double-blind peer review