arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-11-04 至 2025-11-04 共收录 18 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 其他安全 18 篇

2510.08872 2025-11-04 cs.AI cs.GT cs.HC cs.LG cs.MA 81%

GTAlign: Game-Theoretic Alignment of LLM Assistants for Social Welfare

Siqi Zhu, David Zhang, Pedro Cisneros-Velarde, Jiaxuan You

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) VMware Research(VMware研究)

专题命中 其他安全 :alignment(title,abstract);分类 cs.AI、cs.LG

Comments 31 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.19739 2025-11-04 cs.CV 78%

FUSE: Label-Free Image-Event Joint Monocular Depth Estimation via Frequency-Decoupled Alignment and Degradation-Robust Fusion

Pihai Sun, Junjun Jiang, Yuanqi Yao, Youyu Chen, Wenbo Zhao, Kui Jiang, Xianming Liu

机构 * Faculty of Computing, Harbin Institute of Technology(计算机学院,哈尔滨工业大学) Zhengzhou Research Institute, Harbin Institute of Technology(郑州研究院,哈尔滨工业大学)

专题命中 其他安全 :alignment(title,abstract)

Comments [IROS 2025, camera ready version]: 8 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.06164 2025-11-04 cs.LG cs.AI 76%

Model Alignment Search

Satchel Grant

机构 * Stanford University(斯坦福大学)

专题命中 其他安全 :alignment(title);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.14500 2025-11-04 cs.AI cs.MA cs.NE 70%

The Digital Ecosystem of Beliefs: does evolution favour AI over humans?

David M. Bossens, Shanshan Feng, Yew-Soon Ong

机构 * Institute of High Performance Computing (IHPC), Agency for Science, Technology and Research (A*STAR) Centre for Frontier AI Research (CFAR), Agency for Science, Technology and Research (A*STAR)(高性能计算研究所(IHPC)、科技研究局(A*STAR)前沿人工智能研究中心(CFAR)、科技研究局(A*STAR)) School of Computer Science Wuhan University(计算机科学学院 武汉大学)

专题命中 其他安全 :safety(abstract);AI safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.14846 2025-11-04 cs.AI cs.CL cs.LO 62%

Where to Search: Measure the Prior-Structured Search Space of LLM Agents

Zhuo-Yang Song

机构 * School of Physics, Peking University(物理学院,北京大学)

专题命中 其他安全 :safety(abstract);分类 cs.CL、cs.AI

Comments 11 pages, 4 figures, 1 table

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.02184 2025-11-04 stat.ML cs.AI cs.CV cs.LG math.ST stat.TH 62%

Double Descent Meets Out-of-Distribution Detection: Theoretical Insights and Empirical Analysis on the role of model complexity

Mouïn Ben Ammar, David Brellmann, Arturo Mendoza, Antoine Manzanera, Gianni Franchi

机构 * U2IS Lab ENSTA Paris(ENSTA巴黎大学U2IS实验室) Safran Tech(萨弗兰技术)

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG

Comments Accepted at NeurIPS 2025 (Conference on Neural Information Processing Systems)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01258 2025-11-04 cs.CL cs.AI 62%

Measuring Algorithmic Partisanship via Zero-Shot Classification and Its Implications on Political Discourse

Nathan Junzi Chen

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

Comments 19 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00002 2025-11-04 cs.LG cs.AI cs.CV 62%

VRScout: Towards Real-Time, Autonomous Testing of Virtual Reality Games

Yurun Wu, Yousong Sun, Burkhard Wunsche, Jia Wang, Elliott Wen

机构 * School of Computer Science University of Auckland(计算机科学学院 奥克兰大学) School of Advanced Technology Xi'an Jiaotong-Liverpool University(先进科技学院 西交利物浦大学)

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.01380 2025-11-04 cs.CL 57%

Confounding Factors in Relating Model Performance to Morphology

Wessel Poelman, Thomas Bauwens, Miryam de Lhoneux

机构 * NLP, Department of Computer Science, KU Leuven(自然语言处理,计算机科学系,鲁文大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL

Comments EMNLP 2025: Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.01289 2025-11-04 cs.CL 57%

FirstAidQA: A Synthetic Dataset for First Aid and Emergency Response in Low-Connectivity Settings

Saiyma Sittul Muna, Rezwan Islam Salvi, Mushfiqur Rahman Mushfique, Ajwad Abrar

机构 * Islamic University of Technology(伊斯兰技术大学)

专题命中 其他安全 :safety(abstract);分类 cs.CL

Comments Accepted at the 5th Muslims in Machine Learning (MusIML) Workshop, co-located with NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.15049 2025-11-04 cs.RO cs.AI cs.MA 57%

HAD-Gen: Human-like and Diverse Driving Behavior Modeling for Controllable Scenario Generation

Cheng Wang, Lingxin Kong, Massimiliano Tamborski, Stefano V. Albrecht

机构 * School of Engineering and Physical Sciences, Heriot-Watt University(赫瑞斯泰德大学工程与物理科学学院) School of Automation and Software Engineering, Shanxi University(山西大学自动化与软件工程学院) School of Informatics, University of Edinburgh(爱丁堡大学信息学院)

专题命中 其他安全 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.07537 2025-11-04 cs.LG 57%

Accident Impact Prediction based on a deep convolutional and recurrent neural network model

Pouyan Sajadi, Mahya Qorbani, Sobhan Moosavi, Erfan Hassannayebi

机构 * Department of Industrial Engineering, Sharif University of Technology(谢里夫理工大学工业工程系) School of Industrial and System Engineering, Georgia Institute of Technology(佐治亚理工学院工业与系统工程学院) Department of Computer Science and Engineering, Ohio State University(俄亥俄州立大学计算机科学与工程系)

专题命中 其他安全 :safety(abstract);分类 cs.LG

Comments 28 pages, 18 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.18462 2025-11-04 cs.LG cs.CR 57%

MistralBSM: Leveraging Mistral-7B for Vehicular Networks Misbehavior Detection

Wissal Hamhoum, Soumaya Cherkaoui

机构 * Department of Computer and Software Engineering(计算机与软件工程系)

专题命中 其他安全 :safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00033 2025-11-04 cs.RO cs.AI 57%

STRIDER: Navigation via Instruction-Aligned Structural Decision Space Optimization

Diqi He, Xuehao Gao, Hao Li, Junwei Han, Dingwen Zhang

机构 * Northwestern Polytechnical University(西北工业大学) Nanyang Technological University(南洋理工大学) Chongqing University of Posts and Telecommunications(重庆邮电大学)

专题命中 其他安全 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.01645 2025-11-04 cs.CV 50%

Enhancing Diffusion-based Restoration Models via Difficulty-Adaptive Reinforcement Learning with IQA Reward

Xiaogang Xu, Ruihang Chu, Jian Wang, Kun Zhou, Wenjie Shu, Harry Yang, Ser-Nam Lim, Hao Chen, Liang Lin

机构 * The Chinese University of Hong Kong(香港中文大学) Tsinghua University(清华大学) Snap Research Shenzhen University(深圳大学) HKUST(香港科技大学) University of Central Florida(佛罗里达大学) UC Davis(加州大学戴维斯分校) Sun Yat-Sen University(孙中山大学)

专题命中 其他安全 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.27255 2025-11-04 cs.CV 50%

Enhancing Spatio-Temporal Zero-shot Action Recognition with Language-driven Description Attributes

Yehna Kim, Young-Eun Kim, Seong-Whan Lee

机构 * Department of Artificial Intelligence, Korea University(韩国大学人工智能系)

专题命中 其他安全 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.19618 2025-11-04 cs.CV 50%

Pragmatic Heterogeneous Collaborative Perception via Generative Communication Mechanism

Junfei Zhou, Penglin Dai, Quanmin Wei, Bingyi Liu, Xiao Wu, Jianping Wang

机构 * Southwest Jiaotong University(西南交通大学) Engineering Research Center of Sustainable Urban Intelligent Transportation, Ministry of Education, China(可持续城市智能交通工程研究中心,教育部,中国) Wuhan University of Technology(武汉理工大学) City University of Hong Kong(香港城市大学)

专题命中 其他安全 :alignment(abstract)

Comments 26 pages, 10 figures, accepted to NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17359 2025-11-04 cs.IR 50%

MLLM-Driven Semantic Identifier Generation for Generative Cross-Modal Retrieval

Tianyuan Li, Lei Wang, Ahtamjan Ahmat, Yating Yang, Bo Ma, Rui Dong, Bangju Han

专题命中 其他安全 :alignment(abstract)

Comments We plan to revise the methodology and update the experimental analysis before resubmission

详情

展开后加载摘要…

URL PDF HTML 收藏