arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-10-08 至 2025-10-08 共收录 44 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 偏好对齐 4 篇

2504.02725 2025-10-08 cs.CL 88%

SAFER: Advancing Safety Alignment via Efficient Ex-Ante Reasoning

Kehua Feng, Keyan Ding, Yuhao Wang, Menghan Li, Fanjunduo Wei, Xinda Wang, Qiang Zhang, Huajun Chen

机构 * College of Computer Science and Technology, Zhejiang University(浙江大学计算机科学与技术学院) ZJU-Hangzhou Global Scientific and Technological Innovation Center, Zhejiang University(浙江大学Hangzhou全球科技创新中心) Polytechnic Institute, Zhejiang University(浙江大学 polytechnic 院) School of Software Technology, Zhejiang University(浙江大学软件技术学院) ZJU-UIUC Institute, Zhejiang University(浙江大学UIUC研究院)

专题命中 偏好对齐 :alignment(title,abstract);safety(title,abstract);分类 cs.CL

Comments 22 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.05410 2025-10-08 cs.CL cs.LG 84%

Aligning Language Models with Clinical Expertise: DPO for Heart Failure Nursing Documentation in Critical Care

Junyi Fan, Li Sun, Negin Ashrafi, Kamiar Alaei, Maryam Pishgar

机构 * University of Southern California(南加州大学) California State University(加州州立大学)

专题命中 偏好对齐 :DPO(title,abstract);safety(abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.05283 2025-10-08 cs.AI cs.CL cs.CV 76%

Beyond Monolithic Rewards: A Hybrid and Multi-Aspect Reward Optimization for MLLM Alignment

Radha Gulhane, Sathish Reddy Indurthi

机构 * Radha Gulhane(独立研究者) Sathish Reddy Indurthi(独立研究者)

专题命中 偏好对齐 :alignment(title);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21432 2025-10-08 cs.CL cs.AI 62%

Towards Locally Deployable Fine-Tuned Causal Large Language Models for Mode Choice Behaviour

Tareq Alsaleh, Bilal Farooq

机构 * Laboratory of Innovations in Transportation (LiTrans), Toronto Metropolitan University, Canada(交通创新实验室(LiTrans)、多伦多 Metropolitan 大学)

专题命中 偏好对齐 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 安全训练 8 篇

2507.12428 2025-10-08 cs.CL cs.AI cs.LG 85%

Can We Predict Alignment Before Models Finish Thinking? Towards Monitoring Misaligned Reasoning Models

Yik Siu Chan, Zheng-Xin Yong, Stephen H. Bach

机构 * Brown University(布朗大学)

专题命中 安全训练 :alignment(title,abstract);safety(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.05865 2025-10-08 cs.AI cs.CV cs.RO 79%

The Safety Challenge of World Models for Embodied AI Agents: A Review

Lorenzo Baraldi, Zifan Zeng, Chongzhe Zhang, Aradhana Nayak, Hongbo Zhu, Feng Liu, Qunli Zhang, Peng Wang, Shiming Liu, Zheng Hu, Angelo Cangelosi, Lorenzo Baraldi

机构 * University of Pisa(比萨大学) Huawei RAMS Lab(华为RAMS实验室) Technical University of Munich(慕尼黑技术大学) Technical University of Berlin(柏林技术大学) University of Manchester(曼彻斯特大学) University of Modena and Reggio Emilia(莫德纳和雷吉奥艾米利亚大学)

专题命中 安全训练 :safety(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.05156 2025-10-08 cs.SE cs.AI cs.CR 79%

VeriGuard: Enhancing LLM Agent Safety via Verified Code Generation

Lesly Miculicich, Mihir Parmar, Hamid Palangi, Krishnamurthy Dj Dvijotham, Mirko Montanari, Tomas Pfister, Long T. Le

机构 * Google Cloud AI Research(谷歌云人工智能研究) Google DeepMind(谷歌DeepMind)

专题命中 安全训练 :safety(title,abstract);分类 cs.AI

Comments 22 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.19823 2025-10-08 cs.LG cs.AI 79%

Persona Features Control Emergent Misalignment

Miles Wang, Tom Dupré la Tour, Olivia Watkins, Alex Makelov, Ryan A. Chi, Samuel Miserendino, Jeffrey Wang, Achyuta Rajaram, Johannes Heidecke, Tejal Patwardhan, Dan Mossing

机构 * OpenAI

专题命中 安全训练 :alignment(abstract);safety(abstract);AI safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.19056 2025-10-08 cs.CL cs.AI cs.LG 75%

An Embarrassingly Simple Defense Against LLM Abliteration Attacks

Harethah Abu Shairah, Hasan Abed Al Kader Hammoud, Bernard Ghanem, George Turkiyyah

机构 * King Abdullah University of Science and Technology(卡布斯大学)

专题命中 安全训练 :alignment(abstract);safety(abstract);分类 cs.CL、cs.AI、cs.LG

Comments preprint - under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.16366 2025-10-08 cs.CL cs.AI cs.CR cs.LG 75%

A Generative Approach to LLM Harmfulness Mitigation with Red Flag Tokens

David Dobre, Mehrnaz Mofakhami, Sophie Xhonneux, Leo Schwinn, Gauthier Gidel

机构 * Université de Montréal(蒙特利尔大学) Mila(Mila研究所) Technical University of Munich(慕尼黑技术大学) Canada CIFAR AI Chair(加拿大CIFAR人工智能主席)

专题命中 安全训练 :safety(abstract);harmlessness(abstract);分类 cs.CL、cs.AI、cs.LG

Comments 15 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.05194 2025-10-08 q-bio.QM 67%

Reinforcement Learning for Clinical Reasoning: Aligning LLMs with ACR Imaging Appropriateness Criteria

Anni Tziakouri, Filippo Menolascina

专题命中 安全训练 :alignment(abstract);trustworthy(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.05169 2025-10-08 cs.CR cs.AI 57%

From Poisoned to Aware: Fostering Backdoor Self-Awareness in LLMs

Guangyu Shen, Siyuan Cheng, Xiangzhe Xu, Yuan Zhou, Hanxi Guo, Zhuo Zhang, Xiangyu Zhang

机构 * Purdue University(普渡大学) Columbia University(哥伦比亚大学)

专题命中 安全训练 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 提示注入 1 篇

2510.05442 2025-10-08 cs.LG cs.AI cs.CL 85%

Adversarial Reinforcement Learning for Large Language Model Agent Safety

Zizhao Wang, Dingcheng Li, Vaishakh Keshava, Phillip Wallis, Ananth Balashankar, Peter Stone, Lukas Rutishauser

机构 * Google(谷歌) Google Deepmind(谷歌DeepMind) The University of Texas at Austin(德克萨斯大学奥斯汀分校) Sony AI(索尼人工智能)

专题命中 提示注入 :safety(title,abstract);prompt injection(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 幻觉与事实性 5 篇

2501.19252 2025-10-08 cs.CV 78%

Inference-Time Text-to-Video Alignment with Diffusion Latent Beam Search

Yuta Oshima, Masahiro Suzuki, Yutaka Matsuo, Hiroki Furuta

机构 * The University of Tokyo(东京大学) Google DeepMind(谷歌DeepMind)

专题命中 幻觉与事实性 :alignment(title,abstract)

Comments Accepted to NeurIPS2025. Website: https://sites.google.com/view/t2v-dlbs and Code: https://github.com/shim0114/T2V-Diffusion-Search

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.12324 2025-10-08 cs.CL cs.AI 62%

Cross-Document Cross-Lingual NLI via RST-Enhanced Graph Fusion and Interpretability Prediction

Mengying Yuan, Wenhao Wang, Zixuan Wang, Yujie Huang, Kangli Wei, Fei Li, Chong Teng, Donghong Ji

机构 * Key Laboratory of Aerospace Information Security and Trusted Computing, Ministry of Education, School of Cyber Science and Engineering, Wuhan University(航空信息安全与可信计算重点实验室,教育部,网络安全与工程学院,武汉大学) Zhejiang University(浙江大学)

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.CL、cs.AI

Comments EMNLP 2025 Main (Camera Ready)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.06125 2025-10-08 cs.LG 57%

Downsized and Compromised?: Assessing the Faithfulness of Model Compression

Moumita Kamal, Douglas A. Talbert

专题命中 幻觉与事实性 :trustworthy(abstract);分类 cs.LG

Comments Submitted to and under review at Springer Machine Learning Journal

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.06070 2025-10-08 cs.CV 50%

There is More to Attention: Statistical Filtering Enhances Explanations in Vision Transformers

Meghna P Ayyar, Jenny Benois-Pineau, Akka Zemmari

机构 * LaBRI, CNRS, Univ. Bordeaux, UMR 5800(LaBRI、CNRS、波尔多大学、UMR 5800)

专题命中 幻觉与事实性 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.05374 2025-10-08 eess.SY cs.SY 50%

Digital Twins for Intelligent Intersections: A Literature Review

Alben Rome Bagabaldo, Jürgen Hackl

专题命中 幻觉与事实性 :safety(abstract)

Comments 29 pages, 2 figures, under review at Transportation Research Interdisciplinary Perspectives journal

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 安全评测 16 篇

2503.10663 2025-10-08 q-bio.NC cs.AI cs.CV cs.LG 81%

Optimal Transport for Brain-Image Alignment: Unveiling Redundancy and Synergy in Neural Information Processing

Yang Xiao, Wang Lu, Jie Ji, Ruimeng Ye, Gen Li, Xiaolong Ma, Bo Hui

机构 * University of Tulsa(图拉大学) Tsinghua University(清华大学) Clemson University(克莱姆森大学) The University of Arizona(亚利桑那大学)

专题命中 安全评测 :alignment(title,abstract);分类 cs.AI、cs.LG

Comments 14pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.06093 2025-10-08 cs.AI 79%

Classical AI vs. LLMs for Decision-Maker Alignment in Health Insurance Choices

Mallika Mainali, Harsha Sureshbabu, Anik Sen, Christopher B. Rauch, Noah D. Reifsnyder, John Meyer, J. T. Turner, Michael W. Floyd, Matthew Molineaux, Rosina O. Weber

机构 * Information Science, Drexel University, Philadelphia, PA 19104 USA Parallax Advanced Research, 4035 Colonel Glenn Hwy, Beavercreek, OH 45431 USA Knexus Research, 174 Waterfront Street, Suite 310, National Harbor, Oxon Hill, MD 20745 USA Information Science \& Computer Science, Drexel University, Philadelphia, PA 19104 USA

专题命中 安全评测 :alignment(title,abstract);分类 cs.AI

Comments 15 pages, 3 figures. Accepted at the Twelfth Annual Conference on Advances in Cognitive Systems (ACS 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23058 2025-10-08 cs.AI cs.LG 73%

Risk Profiling and Modulation for LLMs

Yikai Wang, Xiaocheng Li, Guanting Chen

机构 * Department of Statistics and Operations Research, UNC-Chapel Hill(统计与运筹学系,北卡罗来纳大学 Chapel Hill 分校) Imperial College Business School, Imperial College London(帝国理工学院伦敦校区商学院)

专题命中 安全评测 :alignment(abstract);RLHF(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00192 2025-10-08 cs.CV 71%

Safe-LLaVA: A Privacy-Preserving Vision-Language Dataset and Benchmark for Biometric Safety

Younggun Kim, Sirnam Swetha, Fazil Kagdi, Mubarak Shah

机构 * Center For Research in Computer Vision, University of Central Florida, USA(计算机视觉研究中心,中央佛罗里达大学) Department of Civil Environmental and Construction Engineering, University of Central Florida, USA(土木环境与建设工程系,中央佛罗里达大学) Department of Computer Science, University of Central Florida, USA(计算机科学系,中央佛罗里达大学)

专题命中 安全评测 :safety(title)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.05976 2025-10-08 cs.CV cs.AI cs.LG 62%

Diffusion Models for Low-Light Image Enhancement: A Multi-Perspective Taxonomy and Performance Analysis

Eashan Adhikarla, Yixin Liu, Brian D. Davison

机构 * Lehigh University(莱维大学)

专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.05972 2025-10-08 cs.CL cs.AI 62%

LexiCon: a Benchmark for Planning under Temporal Constraints in Natural Language

Periklis Mantenoglou, Rishi Hazra, Pedro Zuidberg Dos Martires, Luc De Raedt

专题命中 安全评测 :safety(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.05310 2025-10-08 cs.CL cs.AI 62%

RAG Makes Guardrails Unsafe? Investigating Robustness of Guardrails under RAG-style Contexts

Yining She, Daniel W. Peterson, Marianne Menglin Liu, Vikas Upadhyay, Mohammad Hossein Chaghazardi, Eunsuk Kang, Dan Roth

机构 * Carnegie Mellon University(卡内基梅隆大学) Oracle Cloud Infrastructure(Oracle 云基础设施) University of Pennsylvania(宾夕法尼亚大学)

专题命中 安全评测 :safety(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.12543 2025-10-08 cs.AI cs.CV cs.LG 62%

Human + AI for Accelerating Ad Localization Evaluation

Harshit Rajgarhia, Shivali Dalmia, Mengyang Zhao, Mukherji Abhishek, Kiran Ganesh

机构 * Centific Global Solutions Inc.(Centific全球解决方案公司)

专题命中 安全评测 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.11676 2025-10-08 cs.LG cs.AI stat.ME stat.ML 62%

SKADA-Bench: Benchmarking Unsupervised Domain Adaptation Methods with Realistic Validation On Diverse Modalities

Yanis Lalou, Théo Gnassounou, Antoine Collas, Antoine de Mathelin, Oleksii Kachaiev, Ambroise Odonnat, Alexandre Gramfort, Thomas Moreau, Rémi Flamary

机构 * École Polytechnique, IP Paris, CMAP, UMR 7641(巴黎理工学院) Université Paris-Saclay, Inria, CEA(巴黎萨克雷大学) Inria(法国国家信息与自动化技术研究院) CEA(法国原子能机构) Centre Borelli, ENS Paris-Saclay(巴黎-萨克雷大学博雷利中心) Università degli Studi di Genova(热那亚大学) Inria, Univ. Rennes 2, CNRS, IRISA(法国国家信息与自动化技术研究院、里昂二大学、法国国家科学研究中心、IRISA)

专题命中 安全评测 :alignment(abstract);分类 cs.AI、cs.LG

Comments Published in Transactions on Machine Learning Research

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03413 2025-10-08 cs.CE cs.AI 57%

Report of the 2025 Workshop on Next-Generation Ecosystems for Scientific Computing: Harnessing Community, Software, and AI for Cross-Disciplinary Team Science

Lois Curfman McInnes, Dorian Arnold, Prasanna Balaprakash, Mike Bernhardt, Beth Cerny, Anshu Dubey, Roscoe Giles, Denice Ward Hood, Mary Ann Leung, Vanessa Lopez-Marrero, Paul Messina, Olivia B. Newton, Chris Oehmen, Stefan M. Wild, Jim Willenbring, Lou Woodley, Tony Baylis, David E. Bernholdt, Chris Camano, Johannah Cohoon, Charles Ferenbaugh, Stephen M. Fiore, Sandra Gesing, Diego Gomez-Zara, James Howison, Tanzima Islam, David Kepczynski, Charles Lively, Harshitha Menon, Bronson Messer, Marieme Ngom, Umesh Paliath, Michael E. Papka, Irene Qualters, Elaine M. Raybourn, Katherine Riley, Paulina Rodriguez, Damian Rouson, Michelle Schwalbe, Sudip K. Seal, Ozge Surer, Valerie Taylor, Lingfei Wu

机构 * Argonne National Laboratory(阿贡国家实验室) Emory University(埃默里大学) Oak Ridge National Laboratory(橡树岭国家实验室) Team Libra(团队Libra) Boston University(波士顿大学) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Sustainable Horizons Institute(可持续远景研究所) Stony Brook University(石溪大学) University of Montana(蒙大拿大学) Pacific Northwest National Laboratory(太平洋西北国家实验室) Lawrence Berkeley National Laboratory(伯克利国家实验室) Sandia National Laboratories(桑塔那国家实验室) Center for Scientific Collaboration and Community Engagement(科学协作与社区参与中心) Lawrence Livermore National Laboratory(劳伦斯利弗莫尔国家实验室) Californi(加利福尼亚)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

Comments 38 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.05519 2025-10-08 cs.CY 57%

Assessing Human Rights Risks in AI: A Framework for Model Evaluation

Vyoma Raman, Camille Chabot, Betsy Popken

专题命中 安全评测 :safety(abstract);分类 cs.CY

Comments AAAI/ACM Conference on AI, Ethics, and Society (AIES) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.05185 2025-10-08 cs.MA cs.CE cs.CY cs.NE cs.SI 57%

AgentZero++: Modeling Fear-Based Behavior

Vrinda Malhotra, Jiaman Li, Nandini Pisupati

专题命中 安全评测 :alignment(abstract);分类 cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏