arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

开放权重模型的行为重编程:认知可塑性与对齐边界

Behavioral Reprogramming of Open-Weights Models: Cognitive Plasticity and Alignment Bounds

Lucia Malíčková

arXiv 2608.13069首次发表:更新:

发表机构

National Supercomputing Centre, Slovakia(斯洛伐克国家超级计算中心)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究通过大规模并行超参数搜索等方法,对开放权重LLMs进行行为重编程,实现了主动苏格拉底式对话框架,明确了PEFT的边界等关键结论,为跨语言行为修改提供了实证框架。

AI 中文摘要

大型语言模型(LLMs)大多被对齐为被动、讨好型助手。我们通过实证评估开放权重架构在严格行为重编程下的认知可塑性,挑战这一默认范式。我们的目标是在严格受限的高性能计算(HPC)条件下,诱导出以高频提问为特征的主动苏格拉底式对话框架。通过包含405个HPC作业的大规模并行超参数搜索,我们为参数高效微调(PEFT)定义了精确的数学边界。我们在LoRA秩r=16处识别出一个架构阈值,并通过大量轮次消融实验证明,根据数据集密度,泛化能力在e∈[2,3]的优化训练窗口内严格达到最优收敛(最小验证损失为0.919)。此外,将模型规模扩展至14B参数时,局部评估困惑度更低(1.414)。后续的直接偏好优化(DPO)成功将潜在的主动行为与局部语法解耦,而严格的跨语言压力测试揭示了零样本角色迁移的能力与结构边界,在亲缘关系较近的语系中表现出稳健的对齐,同时在形态差异较大的目标中存在可识别的退化路径。这些发现为计算高效的跨语言行为修改建立了严谨的实证框架。

英文摘要

Large language models (LLMs) are predominantly aligned to function as passive, sycophantic assistants. We challenge this default paradigm by empirically evaluating the cognitive plasticity of open-weight architectures when subjected to rigorous behavioral reprogramming. Our objective is to induce a proactive, Socratic conversational framework, characterized by high-frequency question generation under strictly constrained high-performance computing (HPC) conditions. Through a massively parallelized hyperparameter sweep comprising 405 HPC jobs, we define precise mathematical bounds for parameter-efficient fine-tuning (PEFT). We identify an architectural threshold at LoRA rank $r=16$ and demonstrate via extensive epoch ablation that generalization capacity strictly reaches its optimal convergence within an optimized training window of $e \in [2, 3]$ depending on dataset density (minimum validation loss of 0.919). Furthermore, scaling model capacity to 14B parameters yielded a lower localized evaluation perplexity (1.414). Subsequent Direct Preference Optimization (DPO) successfully decoupled the underlying assertive behavior from localized syntax, while rigorous cross-lingual stress testing reveals both the capabilities and the structural boundaries of zero-shot persona transfer, demonstrating robust alignment in closely related linguistic families alongside identifiable degradation pathways in morphologically distant targets. These findings establish a rigorous empirical framework for compute-efficient, cross-lingual behavioral modification.

CommentsPreprint submitted to arXiv, August 12, 2026. 13 pages, 5 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑