arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2026-01-07 至 2026-01-07 共收录 6 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 其他安全 6 篇

2601.03047 2026-01-07 cs.LG 70%

When the Coffee Feature Activates on Coffins: An Analysis of Feature Extraction and Steering for Mechanistic Interpretability

当咖啡特征在棺材上激活:对特征提取和转向用于机制可解释性的分析

Raphael Ronge, Markus Maier, Frederick Eberhardt

机构 * Department of Philosophy of Nature and Technology(自然哲学与技术系) Munich School of Philosophy(慕尼黑哲学学院) Division of the Humanities and Social Sciences(人文与社会科学系) California Institute of Technology(加州理工学院)

专题命中 其他安全 :safety(abstract);AI safety(abstract);分类 cs.LG

AI总结 本文分析了通过稀疏自编码器提取特征和控制模型输出的方法,指出其在机制可解释性中的局限性和可靠性问题,强调需转向更可靠的预测与控制。

Comments 33 pages (65 with appendix), 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.01584 2026-01-07 cs.CL 57%

Steerability of Instrumental-Convergence Tendencies in LLMs

在大语言模型中操控工具收敛倾向的可行性

Jakub Hoscilowicz

机构 * Warsaw University of Technology(华沙技术大学)

专题命中 其他安全 :safety(abstract);分类 cs.CL

AI总结 该研究通过反工具提示显著降低大语言模型的收敛率,揭示了可控性与安全之间的矛盾,为AI系统的设计提供了新的安全策略。

Comments Code is available at https://github.com/j-hoscilowicz/instrumental_steering

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.03233 2026-01-07 cs.CV 50%

LTX-2: Efficient Joint Audio-Visual Foundation Model

LTX-2:高效的音频-视觉联合基础模型

Yoav HaCohen, Benny Brazowski, Nisan Chiprut, Yaki Bitterman, Andrew Kvochko, Avishai Berkowitz, Daniel Shalem, Daphna Lifschitz, Dudu Moshe, Eitan Porat, Eitan Richardson, Guy Shiran, Itay Chachy, Jonathan Chetboun, Michael Finkelson, Michael Kupchick, Nir Zabari, Nitzan Guetta, Noa Kotler, Ofir Bibi, Ori Gordon, Poriya Panet, Roi Benita, Shahar Armon, Victor Kulikov, Yaron Inger, Yonatan Shiftan, Zeev Melumian, Zeev Farbman

机构 * Lightricks(利光科技)

专题命中 其他安全 :alignment(abstract)

AI总结 LTX-2是一种高效的音频-视觉联合基础模型,通过统一的双流变压器架构生成高质量的音频视觉内容,支持多语言提示和模态感知的可控生成。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.02852 2026-01-07 physics.flu-dyn 50%

ML enhanced measurement of the electrostatic charge distribution of powder conveyed through a duct

基于机器学习的粉末在管道中电荷分布测量

Christoph Wilms, Wenchao Xu, Gizem Ozler, Simon Jantač, Sonja Schmelter, Holger Grosshans

专题命中 其他安全 :safety(abstract)

AI总结 本文提出利用浅层神经网络,通过一维测量数据估计管道中粉体的二维电荷分布,以提高过程安全性。

Journal ref Journal of Loss Prevention in the Process Industries, 92, 105474 (2024)

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.02779 2026-01-07 eess.SY cs.SY 50%

Hierarchical Preemptive Holistic Collaborative Systems for Embodied Multi-Agent Systems: Framework, Hybrid Stability, and Scalability Analysis

分层抢先整体协作系统用于具身多智能体系统:框架、混合稳定性与可扩展性分析

Ting Peng

专题命中 其他安全 :safety(abstract)

AI总结 本文提出分层抢先整体协作框架,通过分解全局协调问题提升具身多智能体系统的安全性和可扩展性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.11160 2026-01-07 eess.AS cs.SD 50%

S2ST-Omni: Hierarchical Language-Aware SpeechLLM Adaptation for Multilingual Speech-to-Speech Translation

S2ST-Omni: 多语言语音到语音翻译的分层语言感知语音LLM适应

Yu Pan, Xiongfei Wu, Yuguang Yang, Jixun Yao, Lei Ma, Jianjun Zhao

机构 * Tencent(腾讯) ByteDance Ltd.(字节跳动)

专题命中 其他安全 :alignment(abstract)

AI总结 S2ST-Omni通过分层语言感知架构和模块化TTS后端,实现多语言语音到语音翻译的高准确性和灵活性。

Comments Working in progress

详情

展开后加载摘要…

URL PDF HTML 收藏