arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

语言模型如何组织与构建道德知识

How Language Models Organize and Structure Moral Knowledge

Orion Reblitz-Richardson

arXiv 2608.27402首次发表:更新:

发表机构

Distiller Labs(蒸馏实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究探究大型语言模型的道德知识组织方式,通过线性探针揭示其道德表征的几何结构,发现其道德整合分量具特异性,且道德困境表征反映冲突结构而非预判决断。

AI 中文摘要

大型语言模型(LLMs)如何组织道德知识?这些模型能广泛检测道德内容,但检测的门槛较低。本文探究它们是否更进一步,区分不同的道德基础并在几何层面组织它们之间的关系。我们在开放权重语言模型上训练六个独立的线性探针,每个探针对应一个道德基础理论(MFT)类别(关怀/伤害、公平/欺骗、自由/压迫、忠诚/背叛、权威/颠覆、神圣/堕落),并在表征空间中研究所得方向之间的关系。我们发现这些方向既不会坍缩为单一的道德检测器,也不会相互孤立,而是跨越了近乎最大数量的独立维度,同时共享一个正的公共分量。该公共分量是整合的标志,相对于以相同方式构建的匹配非道德概念组而言具有道德特异性(平均成对余弦相似度为0.26,而非道德概念组为0.013)。这种几何结构在不同架构和规模之间保持一致,且在预训练早期就达到了整合状态,远早于探针准确率饱和之时。模型发现的结构未显示出道德基础理论所预测的个体化/绑定区分的证据(该检验的效力不足:仅存在20个候选划分),而是反映了语料库统计特征。扩展到道德困境场景,每个困境方向由其组成的道德基础部分构成,其强度是不匹配对基线的2.7倍,同时其大部分方差编码了特定于冲突的结构。模型所表征的是道德张力本身,而非预先解决的判断。

英文摘要

How do large language models (LLMs) organize moral knowledge? Models detect moral content broadly, but detection is a low bar. We ask whether they go further, distinguishing moral foundations from one another and organizing the relationships between them geometrically. We train six independent linear probes on open-weight language models, one per Moral Foundations Theory (MFT) category (care/harm, fair/cheat, lib/oppress, loy/betray, auth/subv, sanc/degrade), and examine how the resulting directions relate to each other in representation space. We find the directions neither collapse into a single moral detector nor isolate from one another. Rather, they span a near-maximal number of independent dimensions while sharing a positive common component. The shared component is the signature of integration, and it is moral-specific relative to a matched non-moral concept battery built identically (mean pairwise cosine 0.26 vs. 0.013). The geometry is consistent across architectures and scale and reaches its integration regime early in pre-training, well before probe accuracy saturates. The structure the model discovers shows no evidence of the individualizing/binding distinction predicted by Moral Foundations Theory (an underpowered test: only 10 distinct splits exist, so it cannot reject at the 0.05 level) but rather reflects corpus statistics. Extending to moral dilemmas, each dilemma direction partially composes from its component foundations, at 2.7x a mismatched-pair baseline, while the majority of its variance encodes conflict-specific structure. The model represents moral tension itself, not a pre-resolved judgment.

Comments32 pages, 16 figures. Code and outputs at https://github.com/deepsteer/deepsteer

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑