arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2026-02-18 至 2026-02-18 共收录 5 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 其他安全 5 篇

2602.15676 2026-02-18 cs.LG cs.AI 81%

Relative Geometry of Neural Forecasters: Linking Accuracy and Alignment in Learned Latent Geometry

神经预报器的相对几何:连接准确性和对齐性在学习潜在几何中的联系

Deniz Kucukahmetler, Maximilian Jean Hemmann, Julian Mosig von Aehrenfeld, Maximilian Amthor, Christian Deubel, Nico Scherf, Diaaeldin Taha

机构 * Max Planck Institute for Human Cognitive and Brain Sciences(马克斯·普朗克人类认知与脑科学研究所) School of Embedded Composite Artificial Intelligence (SECAI)(嵌入式复合人工智能学院) Leipzig University(莱比锡大学) Max Planck Institute for Mathematics in the Sciences(马克斯·普朗克数学研究院)

专题命中 其他安全 :alignment(title,abstract);分类 cs.AI、cs.LG

AI总结 本研究通过相对几何方法探讨神经预报器的对齐与准确性关系,揭示了不同模型家族在表示动态结构上的差异及预测性能的关联。

Comments Accepted to Transactions on Machine Learning Research (TMLR)

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.15368 2026-02-18 cs.CV cs.AI cs.LG eess.IV 81%

GMAIL: Generative Modality Alignment for generated Image Learning

GMAIL: 生成模态对齐用于生成图像学习

Shentong Mo, Sukmin Yun

机构 * Department of Machine Learning, CMU, USA(卡内基梅隆大学机器学习系) Department of Machine Learning, MBZUAI, UAE(马斯克大学人工智能研究所) Department of Artificial Intelligence, Hanyang University ERICA, South Korea(翰阳大学ERICA人工智能系)

专题命中 其他安全 :alignment(title,abstract);分类 cs.AI、cs.LG

AI总结 GMAIL通过多模态学习方法对齐生成图像与真实图像,提升视觉-语言任务中的生成图像学习效果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.15322 2026-02-18 cs.LG cs.AI 62%

On Surprising Effectiveness of Masking Updates in Adaptive Optimizers

自适应优化器中掩码更新的意外有效性

Taejong Joo, Wenhan Xia, Cheolmin Kim, Ming Zhang, Eugene Ie

机构 * Northwestern University(西北大学) Google(谷歌)

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

AI总结 Magma通过动量对齐梯度掩码提升自适应优化器性能,显著降低大型语言模型的困惑度。

Comments Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.15600 2026-02-18 cs.SI cs.AI econ.EM stat.AP 57%

The geometry of online conversations and the causal antecedents of conflictual discourse

在线对话的几何结构与冲突性言论的因果前因

Carlo Santagiustina, Caterina Cruciani

专题命中 其他安全 :alignment(abstract);分类 cs.AI

AI总结 本文研究了在线对话中冲突性语言的因果因素,通过分析时间、对话结构等特征,揭示了回复语气和情感框架的趋同现象。

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.00168 2026-02-18 cs.CV 50%

SSL4EO-S12 v1.1: A Multimodal, Multiseasonal Dataset for Pretraining, Updated

SSL4EO-S12 v1.1:一种用于预训练的多模态、多季节数据集,更新版

Benedikt Blumenstiel, Nassim Ait Ali Braham, Conrad M Albrecht, Stefano Maurogiovanni, Paolo Fraccaro

机构 * IBM Research Europe(IBM欧洲研究中心) German Aerospace Center(德国航空航天中心) Julich Supercomputing Centre University of Iceland(朱利奇超级计算中心爱沙尼亚大学)

专题命中 其他安全 :alignment(abstract)

AI总结 SSL4EO-S12 v1.1通过增加多模态数据和改进数据结构,为预训练大规模基础模型提供了更高效、更全面的地球观测数据集。

详情

展开后加载摘要…

URL PDF HTML 收藏