arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2026-01-29 至 2026-01-29 共收录 10 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 其他安全 10 篇

2502.15301 2026-01-29 cond-mat.soft nlin.AO 78%

Collective behaviors of self-propelled particles with tunable alignment angles

具有可调对齐角度的自驱动粒子集体行为

Zichen Qin, Nariya Uchida

专题命中 其他安全 :alignment(title,abstract)

AI总结 本文提出了一种具有可调对齐角度的自驱动粒子模型,通过数值模拟发现其在特定参数范围内表现出反平行带等独特集体行为,揭示了多体相互作用对向列序的破坏作用。

Comments 7 pages, 5 figures

Journal ref Phys. Rev. E 113, 015411 (2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.23919 2026-01-29 cs.IT cs.SY eess.SY math.IT 78%

Intelligent Angle Map-based Beam Alignment for RIS-aided mmWave Communication Networks

基于智能角度图的RIS辅助毫米波通信网络波束对齐

Hao Xia, Qing Xue, Yanping Liu, Binggui Zhou, Meng Hua, Qianbin Chen

专题命中 其他安全 :alignment(title,abstract)

AI总结 本文提出基于智能角度图的波束对齐方法,利用Transformer模型实现高效波束对齐,无需扫描即可实现高精度通信性能。

Journal ref IEEE Transactions on Network Science and Engineering, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.10138 2026-01-29 cs.LG 74%

Provably Efficient RL under Episode-Wise Safety in Constrained MDPs with Linear Function Approximation

在具有线性函数逼近的约束马尔可夫决策过程中的episode级安全强化学习的可证明效率

Toshinori Kitamura, Arnob Ghosh, Tadashi Kozuno, Wataru Kumagai, Kazumi Kasaura, Kenta Hoshino, Yohei Hosoe, Yutaka Matsuo

专题命中 其他安全 :safety(title);分类 cs.LG

AI总结 本文提出了一种在约束马尔可夫决策过程中的强化学习算法,实现episode级安全性和$\tilde{\mathcal{O}}(\sqrt{K})$的遗憾,同时具有多项式计算效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.11558 2026-01-29 cs.CV cs.AI cs.CL 62%

DaMO: A Data-Efficient Multimodal Orchestrator for Temporal Reasoning with Video LLMs

DaMO:一种数据高效的多模态协调器,用于视频LLM的时序推理

Bo-Cheng Chiu, Jen-Jee Chen, Yu-Chee Tseng, Feng-Chi Chen, An-Zi Yen

机构 * College of Artificial Intelligence, National Yang Ming Chiao Tung University(人工智能学院,阳明交通大学) Institute of Population Health Sciences, National Health Research Institutes(人口健康科学研究所,国家健康研究院) Department of Computer Science, National Yang Ming Chiao Tung University(计算机科学系,阳明交通大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

AI总结 DaMO是一种专为视频LLM设计的数据高效多模态协调器,通过时序感知Fuseformer和四阶段训练范式提升时序推理能力,实现更精确的多模态理解。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.20467 2026-01-29 cs.AI cs.CL 62%

CtrlCoT: Dual-Granularity Chain-of-Thought Compression for Controllable Reasoning

CtrlCoT:用于可控推理的双粒度推理链压缩

Zhenxuan Fan, Jie Cao, Yang Dai, Zheqi Lv, Wenqiao Zhang, Zhongle Xie, Peng LU, Beng Chin Ooi

机构 * Zhejiang University(浙江大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

AI总结 CtrlCoT通过双粒度方法实现高效可控推理,结合语义抽象与标记级删除,提升推理效率和可靠性。

Comments 16 pages, 9 figures, 11 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.14174 2026-01-29 cs.LG cs.AI 62%

Physics-Guided Multimodal Transformers are the Necessary Foundation for the Next Generation of Meteorological Science

物理引导的多模态Transformer是下一代气象科学的必要基础

Jing Han, Hanting Chen, Kai Han, Xiaomeng Huang, Wenjun Xu, Dacheng Tao, Ping Zhang

机构 * School of Artificial Intelligence, Beijing University of Posts Huawei Noah's Ark Lab Department of Earth System Science, Tsinghua University College of Computing \& Data Science, Nanyang Technological University State Key Laboratory of Networking Switching Technology, Beijing University of Posts

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

AI总结 本文提出通过物理引导的多模态Transformer构建下一代气象科学的统一范式,以提升模型的科学一致性和物理约束能力。

Comments Perspective article

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.20747 2026-01-29 cs.CL cs.HC 57%

Like a Therapist, But Not: Reddit Narratives of AI in Mental Health Contexts

像治疗师,但不是:Reddit中关于心理健康情境下AI的叙述

Elham Aghakhani, Rezvaneh Rezapour

机构 * Drexel University(德雷塞尔大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL

AI总结 本研究通过分析Reddit上的心理健康相关AI使用叙述,探讨了用户对AI在心理健康领域中的评价与关系契合,揭示了任务契合度与情感联系在用户互动中的不同影响。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.02072 2026-01-29 cs.LG cs.IR 57%

Abex-rat: Synergizing Abstractive Augmentation and Adversarial Training for Classification of Occupational Accident Reports

ABEX-RAT:协同抽象增强与对抗训练用于职业事故报告分类

Jian Chen, Jiabao Dou

机构 * Ningxia Research Institute of Transport Science(宁夏交通科学研究院) Department of Computer Science(计算机科学系)

专题命中 其他安全 :safety(abstract);分类 cs.LG

AI总结 ABEX-RAT通过结合抽象增强与对抗训练,高效解决职业事故报告分类中的类别不平衡问题,实现90.32%的宏F1得分。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.20075 2026-01-29 cs.CV 50%

Sparse CLIP: Co-Optimizing Interpretability and Performance in Contrastive Learning

稀疏CLIP:在对比学习中协同优化可解释性与性能

Chuan Qin, Constantin Venhoff, Sonia Joseph, Fanyi Xiao, Stefan Scherer

专题命中 其他安全 :alignment(abstract)

AI总结 稀疏CLIP通过在对比学习中直接整合稀疏性,实现了可解释性与性能的协同优化,保留了多模态能力并提升了下游任务性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.22134 2026-01-29 cs.HC 50%

IntentFlow: Investigating Fluid Dynamics of Intent Communication in Generative AI

IntentFlow:探究生成AI中意图通信的流体动力学

Yoonsu Kim, Kihoon Son, Seoyoung Kim, Brandon Chin, Juho Kim

专题命中 其他安全 :alignment(abstract)

AI总结 IntentFlow通过系统分析揭示生成AI中意图通信的关键要素及交互机制,提出四种核心支持方面并验证其在写作任务中的有效性。

详情

展开后加载摘要…

URL PDF HTML 收藏