arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 8057 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 其他安全 8057 篇

2511.07086 2025-11-11 cs.AI 57%

LLM Driven Processes to Foster Explainable AI

Marcel Pehlke, Marc Jansen

机构 * University of Applied Sciences Ruhr West(鲁尔西部应用科学大学)

专题命中 其他安全 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.06552 2025-11-11 cs.SE cs.AI 57%

LLM For Loop Invariant Generation and Fixing: How Far Are We?

Mostafijur Rahman Akhond, Saikat Chakraborty, Gias Uddin

机构 * York University(约克大学) Microsoft Research(微软研究院)

专题命中 其他安全 :safety(abstract);分类 cs.AI

Comments This work has been submitted to the IEEE for possible publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.15052 2025-11-11 cs.AI 57%

GlitchMiner: Mining Glitch Tokens in Large Language Models via Gradient-based Discrete Optimization

Zihui Wu, Haichang Gao, Ping Wang, Shudong Zhang, Zhaoxiang Liu, Shiguo Lian

专题命中 其他安全 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.05769 2025-11-11 cs.HC cs.AI 57%

Lived Experience in Dialogue: Co-designing Personalization in Large Language Models to Support Youth Mental Well-being

Kathleen W. Guan, Sarthak Giri, Mohammed Amara, Bernard J. Jansen, Enrico Liscio, Milena Esherick, Mohammed Al Owayyed, Ausrine Ratkute, Gayane Sedrakyan, Mark de Reuver, Joao Fernando Ferreira Goncalves, Caroline A. Figueroa

专题命中 其他安全 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.05528 2025-11-11 cs.AI 57%

SMAGDi: Socratic Multi Agent Interaction Graph Distillation for Efficient High Accuracy Reasoning

Aayush Aluru, Myra Malik, Samarth Patankar, Spencer Kim, Kevin Zhu, Sean O'Brien, Vasu Sharma

机构 * Algoverse AI Research(Algoverse AI研究)

专题命中 其他安全 :alignment(abstract);分类 cs.AI

Comments Multi-Turn Interactions in Large Language Models (MTI-LLM) Workshop at NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.13003 2025-11-11 cs.CL 57%

OPLoRA: Orthogonal Projection LoRA Prevents Catastrophic Forgetting during Parameter-Efficient Fine-Tuning

Yifeng Xiong, Xiaohui Xie

专题命中 其他安全 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.05325 2025-11-10 cs.LG 57%

Turning Adversaries into Allies: Reversing Typographic Attacks for Multimodal E-Commerce Product Retrieval

Janet Jenq, Hongda Shen

机构 * PitchBook USA(PitchBook美国公司)

专题命中 其他安全 :alignment(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.06980 2025-11-10 cs.CL 57%

Exploring Multimodal Perception in Large Language Models Through Perceptual Strength Ratings

Jonghyun Lee, Dojun Park, Jiwoo Lee, Hoekeon Choi, Sung-Eun Lee

专题命中 其他安全 :alignment(abstract);分类 cs.CL

Comments Published in IEEE Access

Journal ref IEEE Access, vol. 13, pp. 176751-176769, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.04139 2025-11-07 cs.CL cs.SD 57%

CantoASR: Prosody-Aware ASR-LALM Collaboration for Low-Resource Cantonese

Dazhong Chen, Yi-Cheng Lin, Yuchen Huang, Ziwei Gong, Di Jiang, Zeying Xie, Yi R., Fung

机构 * The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)) Hong Kong University of Science and Technology(香港科学与技术大学) National Taiwan University, Taiwan(台湾国立台湾大学) Columbia University(哥伦比亚大学) WeBank Co., Ltd., Shenzhen, China(深圳网商银行有限公司)

专题命中 其他安全 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.03410 2025-11-06 cs.CL 57%

Knowledge-Augmented Question Error Correction for Chinese Question Answer System with QuestionRAG

Longpeng Qiu, Ting Li, Shuai Mao, Nan Yang, Xiaohui Yan

机构 * University of Chinese Academy of Sciences(中国科学院大学) Huawei Technologies Co., Ltd.(华为技术有限公司)

专题命中 其他安全 :alignment(abstract);分类 cs.CL

Comments EMNLP2025 Industry Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15737 2025-11-06 cs.LG cs.CV 57%

Probability Density from Latent Diffusion Models for Out-of-Distribution Detection

Joonas Järve, Karl Kaspar Haavel, Meelis Kull

机构 * Institute of Computer Science, University of Tartu, Estonia(计算机科学研究所,塔尔图大学,爱沙尼亚)

专题命中 其他安全 :safety(abstract);分类 cs.LG

Comments ECAI 2025

Journal ref Frontiers in Artificial Intelligence and Applications 413 (ECAI 2025) 5027 - 5034

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.02690 2025-11-05 cs.LG 57%

Curriculum Design for Trajectory-Constrained Agent: Compressing Chain-of-Thought Tokens in LLMs

Georgios Tzannetos, Parameswaran Kamalaruban, Adish Singla

机构 * MPI-SWS(马克斯·普朗克所际研究所)

专题命中 其他安全 :safety(abstract);分类 cs.LG

Comments NeurIPS'25 paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.02606 2025-11-05 cs.AI cs.HC 57%

A Multi-Agent Psychological Simulation System for Human Behavior Modeling

Xiangen Hu, Jiarui Tong, Sheng Xu

专题命中 其他安全 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21971 2025-11-05 cs.LG 57%

GRAM-DTI: adaptive multimodal representation learning for drug target interaction prediction

Feng Jiang, Amina Mollaysa, Hehuan Ma, Tommaso Mansi, Junzhou Huang, Mangal Prakash, Rui Liao

机构 * University of Texas at Arlington(德克萨斯理工大学) Johnson & Johnson Innovative Medicine(强生创新医学)

专题命中 其他安全 :alignment(abstract);分类 cs.LG

Journal ref NeurIPS 2025 2nd Workshop on Multi-modal Foundation Models and Large Language Models for Life Sciences

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.06910 2025-11-05 cs.CL 57%

Identifying Aspects in Peer Reviews

Sheng Lu, Ilia Kuznetsov, Iryna Gurevych

机构 * Ubiquitous Knowledge Processing Lab (UKP Lab)(通用知识处理实验室) Department of Computer Science(计算机科学系) Hessian Center for AI (hessian.AI)(黑森人工智能中心)

专题命中 其他安全 :alignment(abstract);分类 cs.CL

Comments EMNLP 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.01380 2025-11-04 cs.CL 57%

Confounding Factors in Relating Model Performance to Morphology

Wessel Poelman, Thomas Bauwens, Miryam de Lhoneux

机构 * NLP, Department of Computer Science, KU Leuven(自然语言处理,计算机科学系,鲁文大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL

Comments EMNLP 2025: Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.01289 2025-11-04 cs.CL 57%

FirstAidQA: A Synthetic Dataset for First Aid and Emergency Response in Low-Connectivity Settings

Saiyma Sittul Muna, Rezwan Islam Salvi, Mushfiqur Rahman Mushfique, Ajwad Abrar

机构 * Islamic University of Technology(伊斯兰技术大学)

专题命中 其他安全 :safety(abstract);分类 cs.CL

Comments Accepted at the 5th Muslims in Machine Learning (MusIML) Workshop, co-located with NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.15049 2025-11-04 cs.RO cs.AI cs.MA 57%

HAD-Gen: Human-like and Diverse Driving Behavior Modeling for Controllable Scenario Generation

Cheng Wang, Lingxin Kong, Massimiliano Tamborski, Stefano V. Albrecht

机构 * School of Engineering and Physical Sciences, Heriot-Watt University(赫瑞斯泰德大学工程与物理科学学院) School of Automation and Software Engineering, Shanxi University(山西大学自动化与软件工程学院) School of Informatics, University of Edinburgh(爱丁堡大学信息学院)

专题命中 其他安全 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.07537 2025-11-04 cs.LG 57%

Accident Impact Prediction based on a deep convolutional and recurrent neural network model

Pouyan Sajadi, Mahya Qorbani, Sobhan Moosavi, Erfan Hassannayebi

机构 * Department of Industrial Engineering, Sharif University of Technology(谢里夫理工大学工业工程系) School of Industrial and System Engineering, Georgia Institute of Technology(佐治亚理工学院工业与系统工程学院) Department of Computer Science and Engineering, Ohio State University(俄亥俄州立大学计算机科学与工程系)

专题命中 其他安全 :safety(abstract);分类 cs.LG

Comments 28 pages, 18 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.18462 2025-11-04 cs.LG cs.CR 57%

MistralBSM: Leveraging Mistral-7B for Vehicular Networks Misbehavior Detection

Wissal Hamhoum, Soumaya Cherkaoui

机构 * Department of Computer and Software Engineering(计算机与软件工程系)

专题命中 其他安全 :safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00033 2025-11-04 cs.RO cs.AI 57%

STRIDER: Navigation via Instruction-Aligned Structural Decision Space Optimization

Diqi He, Xuehao Gao, Hao Li, Junwei Han, Dingwen Zhang

机构 * Northwestern Polytechnical University(西北工业大学) Nanyang Technological University(南洋理工大学) Chongqing University of Posts and Telecommunications(重庆邮电大学)

专题命中 其他安全 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.27432 2025-11-03 cs.CV cs.AI 57%

Mitigating Semantic Collapse in Partially Relevant Video Retrieval

WonJun Moon, MinSeok Jung, Gilhan Park, Tae-Young Kim, Cheol-Ho Cho, Woojin Jun, Jae-Pil Heo

专题命中 其他安全 :alignment(abstract);分类 cs.AI

Comments Accpeted to NeurIPS 2025. Code is available at https://github.com/admins97/MSC_PRVR

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.09391 2025-10-31 cs.CL 57%

Comparing human and LLM politeness strategies in free production

Haoran Zhao, Robert D. Hawkins

机构 * Department of Linguistics University of Washington(语言学系华盛顿大学) Department of Linguistics Stanford University(语言学系斯坦福大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL

Comments 25 pages, 5 figures | EMNLP 2025 camera-ready version

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.26242 2025-10-31 cs.AI 57%

Retrieval Augmented Generation-Enhanced Distributed LLM Agents for Generalizable Traffic Signal Control with Emergency Vehicles

Xinhang Li, Qing Guo, Junyu Chen, Zheng Guo, Shengzhe Xu, Lei Li, Lin Zhang

机构 * School of Artificial Intelligence, Beijing University of Posts(人工智能学院,北京邮电大学) Beijing Big Data Center(北京大数据中心)

专题命中 其他安全 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23127 2025-10-31 cs.AI 57%

Lost in Tokenization: Context as the Key to Unlocking Biomolecular Understanding in Scientific LLMs

Kai Zhuang, Jiawei Zhang, Yumou Liu, Hanqun Cao, Chunbin Gu, Mengdi Liu, Zhangyang Gao, Zitong Jerry Wang, Xuanhe Zhou, Pheng-Ann Heng, Lijun Wu, Conghui He, Cheng Tan

机构 * Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Westlake University(西交大学) Shanghai Innovation Institute(上海创新研究院) Shanghai Jiaotong University(上海交通大学) The Chinese University of Hong Kong(香港中文大学) Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所)

专题命中 其他安全 :alignment(abstract);分类 cs.AI

Comments 38 pages, under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16629 2025-10-31 cs.LG 57%

On the Impossibility of Retrain Equivalence in Machine Unlearning

Jiatong Yu, Yinghui He, Anirudh Goyal, Sanjeev Arora

机构 * Princeton Language and Intelligence(普林斯顿语言与智能)

专题命中 其他安全 :alignment(abstract);分类 cs.LG

Comments Code available at https://princeton-pli.github.io/impossibility-unlearning/

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.15811 2025-10-31 cs.LG 57%

On the creation of narrow AI: hierarchy and nonlocality of neural network skills

Eric J. Michaud, Asher Parker-Sartori, Max Tegmark

机构 * Department of Physics, Massachusetts Institute of Technology(物理学系,麻省理工学院) Department of EECS, Massachusetts Institute of Technology(电子工程与计算机科学系,麻省理工学院) The NSF AI Institute for Artificial Intelligence and Fundamental Interactions(国家科学基金会人工智能与基本相互作用研究所)

专题命中 其他安全 :safety(abstract);分类 cs.LG

Comments NeurIPS 2025; 20 pages, 13 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.25150 2025-10-30 cs.CL 57%

Explainable Disentanglement on Discrete Speech Representations for Noise-Robust ASR

Shreyas Gopal, Ashutosh Anshul, Haoyang Li, Yue Heng Yeo, Hexin Liu, Eng Siong Chng

机构 * College of Computing and Data Science, Nanyang Technological University(计算与数据科学学院,南洋理工大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL

Comments Awarded Best Student Paper at APSIPA ASC 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23822 2025-10-30 cs.AI 57%

ReCAP: Recursive Context-Aware Reasoning and Planning for Large Language Model Agents

Zhenyu Zhang, Tianyi Chen, Weiran Xu, Alex Pentland, Jiaxin Pei

机构 * Department of Computer Science, Stanford University(斯坦福大学计算机科学系) Stanford Institute for Human-Centered AI(斯坦福大学人本人工智能研究所) MIT Media Lab(麻省理工学院媒体实验室)

专题命中 其他安全 :alignment(abstract);分类 cs.AI

Journal ref 39th Conference on Neural Information Processing Systems (NeurIPS 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.24331 2025-10-29 cs.LG cs.CV 57%

What do vision-language models see in the context? Investigating multimodal in-context learning

Gabriel O. dos Santos, Esther Colombini, Sandra Avila

机构 * Instituto de Computação, Universidade Estadual de Campinas (UNICAMP)(计算机学院,Campinas州立大学(UNICAMP))

专题命中 其他安全 :alignment(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏