发表机构
The George Washington University(乔治·华盛顿大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究类ChatGPT人工智能出现不良内容转向的原因,通过多体相互作用等分析表明此类故障是可预见工程风险,提出基于多体相互作用的分析方法及少数盆地简化等,得出对人工智能危害评估有重要意义的结论。
AI 中文摘要
为何类ChatGPT人工智能尽管在架构和训练上存在重大差异,即使在确定性贪婪解码下也会意外地转向不良内容(如有害、误导、重复)?我们表明,这类广泛的翻转是由令牌(自旋)在穿越有限层系统时的多体相互作用引起的。翻转作为竞争输出盆地之间的动态首次通过过程出现。注意力紊乱控制着朝向、远离或沿着盆地边界的传输。少数盆地简化产生一个封闭的有限层阈值,其粗粒度预测在类ChatGPT家族中显示出良好的一致性。这些结果表明,一类广泛的人工智能故障代表了“可预见的工程风险”,而非固有不可预测的行为,这对人工智能危害的法律和社会评估具有重要意义。
英文摘要
Why do ChatGPT-like AIs, despite major architectural and training differences, unexpectedly tip to undesirable content (e.g. harmful, misleading, repetitive) even under deterministic greedy decoding? We show that a broad class of such tippings is caused by the many-body interactions between tokens (spins) as they cross the finite-layer system. Tipping emerges as a dynamical first passage process between competing output basins. Attention disorder controls the transport toward, away from, or along the basins' boundary. A few-basin reduction yields a closed finite-layer threshold, whose coarse-grained predictions show good agreement across ChatGPT-like families. These results suggest that a broad class of AI failures represents 'foreseeable engineering risk' rather than inherently unpredictable behavior, with important implications for legal and societal assessments of AI harm.