arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-09-25 至 2025-09-25 共收录 37 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 偏好对齐 2 篇

2504.08590 2025-09-25 cs.CL 70%

Playpen: An Environment for Exploring Learning Through Conversational Interaction

Nicola Horst, Davide Mazzaccara, Antonia Schmidt, Michael Sullivan, Filippo Momentè, Luca Franceschetti, Philipp Sadler, Sherzod Hakimov, Alberto Testoni, Raffaella Bernardi, Raquel Fernández, Alexander Koller, Oliver Lemon, David Schlangen, Mario Giulianelli, Alessandro Suglia

机构 * University of Potsdam(波恩大学) University of Trento(特伦特大学) Saarland University(萨尔大学) ETH Zurich(苏黎世联邦理工学院) Amsterdam UMC(阿姆斯特丹医疗中心) Free University of Bozen Bolzano(博兹纳自由大学) University of Amsterdam(阿姆斯特丹大学) Heriot-Watt University(赫里奥特-瓦特大学) DFKI(德意志联邦研究所) UCL(伦敦大学学院) University of Edinburgh(爱丁堡大学)

专题命中 偏好对齐 :alignment(abstract);DPO(abstract);分类 cs.CL

Comments Accepted at EMNLP 2025 (Main) Source code: https://github.com/lm-playpen/playpen Please send correspodence to: lm-playschool@googlegroups.com

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19336 2025-09-25 cs.CL cs.AI 62%

Cognitive-Level Adaptive Generation via Capability-Aware Retrieval and Style Adaptation

Qingsong Wang, Tao Wu, Wang Lin, Yueying Feng, Gongsheng Yuan, Chang Yao, Jingyuan Chen

机构 * Zhejiang University(浙江大学)

专题命中 偏好对齐 :alignment(abstract);分类 cs.CL、cs.AI

Comments Accepted to Findings of EMNLP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 安全训练 3 篇

2509.19688 2025-09-25 cs.RO cs.LG cs.SY eess.SY math.OC 79%

Formal Safety Verification and Refinement for Generative Motion Planners via Certified Local Stabilization

Devesh Nath, Haoran Yin, Glen Chou

机构 * Georgia Institute of Technology(佐治亚理工学院)

专题命中 安全训练 :safety(title,abstract);分类 cs.LG

Comments 10 pages, 12 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19839 2025-09-25 cs.AI 70%

LatentGuard: Controllable Latent Steering for Robust Refusal of Attacks and Reliable Response Generation

Huizhen Shu, Xuying Li, Zhuo Li

专题命中 安全训练 :alignment(abstract);safety(abstract);分类 cs.AI

Comments 9-page NeurIPS 2025 preprint including 3 figures and 1 table, with additional appendix material. Prepared using the NeurIPS 2025 preprint template and compiled with pdfLaTeX. All references are included via the provided .bbl file. Figures are in PDF format. No external supplementary files. All necessary style files and images are included

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.12621 2025-09-25 cs.CL cs.IR 70%

SAFE: Improving LLM Systems using Sentence-Level In-generation Attribution

João Eduardo Batista, Emil Vatai, Mohamed Wahib

机构 * RIKEN-CCS Kobe, Japan(日本神户RIKEN-CCS)

专题命中 安全训练 :safety(abstract);trustworthy(abstract);分类 cs.CL

Comments 30 pages (9 pages of content, 5 pages of references, 16 pages of supplementary material), 7 figures, 13 tables

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 越狱攻击 2 篇

2509.19775 2025-09-25 cs.CL cs.AI cs.CR 86%

bi-GRPO: Bidirectional Optimization for Jailbreak Backdoor Injection on LLMs

Wence Ji, Jiancan Wu, Aiying Li, Shuyi Zhang, Junkang Wu, An Zhang, Xiang Wang, Xiangnan He

机构 * University of Science and Technology of China(中国科学技术大学)

专题命中 越狱攻击 :jailbreak(title,abstract);RLHF(abstract);safety(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.13739 2025-09-25 cs.CV 67%

Enhancing Targeted Adversarial Attacks on Large Vision-Language Models via Intermediate Projector

Yiming Cao, Yanjie Li, Kaisheng Liang, Bin Xiao

专题命中 越狱攻击 :alignment(abstract);safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 幻觉与事实性 1 篇

2509.19375 2025-09-25 cs.LG cs.AI stat.ML 62%

Uncertainty Quantification of Large Language Models using Approximate Bayesian Computation

Mridul Sharma, Adeetya Patel, Zaneta D' Souza, Samira Abbasgholizadeh Rahimi, Siva Reddy, Sreenath Madathil

机构 * Faculty of Dental Medicine and Oral Health Sciences, McGill University(牙医学院与口腔健康科学学院,麦吉尔大学) McGill University(麦吉尔大学) Mila–Quebec Artificial Intelligence Institute(魁北克人工智能研究所) School of Computer Science, McGill University(计算机科学学院,麦吉尔大学)

专题命中 幻觉与事实性 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 隐私与版权 1 篇

2410.15894 2025-09-25 cs.OS 50%

MVVM: Deploy Your AI Agents-Securely, Efficiently, Everywhere

Yiwei Yang, Aibo Hu, Yusheng Zheng, Brian Zhao, Xinqi Zhang, Dawei Xiang, Kexin Chu, Wei Zhang, Andi Quinn

专题命中 隐私与版权 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

6. 安全评测 12 篇

2509.20146 2025-09-25 cs.CV cs.AI 70%

EchoBench: Benchmarking Sycophancy in Medical Large Vision-Language Models

Botai Yuan, Yutian Zhou, Yingjie Wang, Fushuo Huo, Yongcheng Jing, Li Shen, Ying Wei, Zhiqi Shen, Ziwei Liu, Tianwei Zhang, Jie Yang, Dacheng Tao

机构 * Nanyang Technological University(南洋理工大学) Shanghai Jiao Tong University(上海交通大学) Fudan University(复旦大学) Sun Yat-sen University(中山大学) Zhejiang University(浙江大学)

专题命中 安全评测 :safety(abstract);trustworthy(abstract);分类 cs.AI

Comments 29 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19325 2025-09-25 cs.CL 70%

How Much of Your Data Can Suck? Thresholds for Domain Performance and Emergent Misalignment in LLMs

Jian Ouyang, Arman T, Ge Jin

机构 * Invisible Technologies(隐形技术)

专题命中 安全评测 :alignment(abstract);safety(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23745 2025-09-25 cs.CV cs.AI cs.LG 62%

To Trust Or Not To Trust Your Vision-Language Model's Prediction

Hao Dong, Moru Liu, Jian Liang, Eleni Chatzi, Olga Fink

专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.08810 2025-09-25 cs.CR cs.AI 57%

Machine Learning-Based Detection of DDoS Attacks in VANETs for Emergency Vehicle Communication

Bappa Muktar, Vincent Fono, Adama Nouboukpo

机构 * Department of Computer Science University of Quebec in Outaouais (UQO)(计算机科学系魁北克大学 Outaouais 分校)

专题命中 安全评测 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19604 2025-09-25 cs.LG 57%

Improved Therapeutic Antibody Reformatting through Multimodal Machine Learning

Jiayi Xin, Aniruddh Raghu, Nick Bhattacharya, Adam Carr, Melanie Montgomery, Hunter Elliott

机构 * University of Pennsylvania(宾夕法尼亚大学) BigHat Biosciences(BigHat生物技术公司)

专题命中 安全评测 :safety(abstract);分类 cs.LG

Comments NeurIPS 2025 AI4Science Workshop and NeurIPS 2025 Multi-modal Foundation Models and Large Language Models for Life Sciences Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.01301 2025-09-25 cs.CL 57%

Culture is Everywhere: A Call for Intentionally Cultural Evaluation

Juhyun Oh, Inha Cha, Michael Saxon, Hyunseung Lim, Shaily Bhatt, Alice Oh

专题命中 安全评测 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.00935 2025-09-25 cs.CR cs.AI 57%

Measuring Harmfulness of Computer-Using Agents

Aaron Xuxiang Tian, Ruofan Zhang, Janet Tang, Ji Wang, Tianyu Shi, Jiaxin Wen

机构 * Arizona State University(亚利桑那州立大学) University of California, Berkeley(加州大学伯克利分校)

专题命中 安全评测 :safety(abstract);分类 cs.AI

Comments 17 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23368 2025-09-25 cs.CL 57%

Threading the Needle: Reweaving Chain-of-Thought Reasoning to Explain Human Label Variation

Beiduo Chen, Yang Janet Liu, Anna Korhonen, Barbara Plank

专题命中 安全评测 :alignment(abstract);分类 cs.CL

Comments Accepted by EMNLP 2025 Main (Oral), 25 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.16180 2025-09-25 cs.CV cs.CL 57%

Redemption Score: A Multi-Modal Evaluation Framework for Image Captioning via Distributional, Perceptual, and Linguistic Signal Triangulation

Ashim Dahal, Ankit Ghimire, Saydul Akbar Murad, Nick Rahimi

机构 * University of Southern Mississippi(密苏里州南方大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.18485 2025-09-25 q-bio.NC cs.CV 50%

Deciphering Functions of Neurons in Vision-Language Models

Jiaqi Xu, Cuiling Lan, Yan Lu

机构 * University of Science and Technology of China(中国科学技术大学) Microsoft Research Asia(微软亚洲研究院)

专题命中 安全评测 :trustworthy(abstract)

Comments Accepted by the 31st ACM International Conference on Multimedia (ACM MM 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19958 2025-09-25 cs.RO 50%

Generalist Robot Manipulation beyond Action Labeled Data

Alexander Spiridonov, Jan-Nico Zaech, Nikolay Nikolov, Luc Van Gool, Danda Pani Paudel

机构 * ETH Zurich, Switzerland(苏黎世联邦理工学院,瑞士)

专题命中 安全评测 :alignment(abstract)

Comments Accepted at Conference on Robot Learning 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19831 2025-09-25 eess.AS 50%

SCORE: Scaling audio generation using Standardized COmposite REwards

Jaemin Jung, Jaehun Kim, Inkyu Shin, Joon Son Chung

专题命中 安全评测 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

7. AI治理与伦理 1 篇

2505.15074 2025-09-25 cs.CL cs.AI cs.LG 75%

DISCO Balances the Scales: Adaptive Domain- and Difficulty-Aware Reinforcement Learning on Imbalanced Data

Yuhang Zhou, Jing Zhu, Shengyi Qian, Zhuokai Zhao, Xiyao Wang, Xiaoyu Liu, Ming Li, Paiheng Xu, Wei Ai, Furong Huang

专题命中 AI治理与伦理 :alignment(abstract);RLHF(abstract);分类 cs.CL、cs.AI、cs.LG

Comments Accepted by EMNLP 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏

8. 其他安全 15 篇

2509.20093 2025-09-25 cs.RO 82%

Hybrid Safety Verification of Multi-Agent Systems using $ψ$-Weighted CBFs and PAC Guarantees

Venkat Margapuri, Garik Kazanjian, Naren Kosaraju

机构 * Department of Computing Sciences at Villanova University(维拉诺瓦大学计算机科学系)

专题命中 其他安全 :safety(title,abstract);alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19657 2025-09-25 cs.CL cs.AI cs.SI 81%

Large Language Models for Pedestrian Safety: An Application to Predicting Driver Yielding Behavior at Unsignalized Intersections

Yicheng Yang, Zixian Li, Jean Paul Bizimana, Niaz Zafri, Yongfeng Dong, Tianyi Li

专题命中 其他安全 :safety(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19745 2025-09-25 cs.CL cs.SD 79%

PART: Progressive Alignment Representation Training for Multilingual Speech-To-Text with LLMs

Pei Zhang, Andong Chen, Xi Chen, Baosong Yang, Derek F. Wong, Fei Huang

机构 * Tongyi Lab, Alibaba Group(通义实验室,阿里巴巴集团) The Chinese University of Hong Kong(香港中文大学) NLP 2 CT Lab, University of Macau(自然语言处理2CT实验室,澳门大学)

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19329 2025-09-25 cs.CL stat.ME 79%

How Model Size, Temperature, and Prompt Style Affect LLM-Human Assessment Score Alignment

Julie Jung, Max Lu, Sina Chole Benker, Dogus Darici

机构 * Harvard Graduate School of Education(哈佛教育研究生院) Munster University(穆恩斯特大学) Institute of Anatomy and Neurobiology, University of Münster(解剖与神经生物学研究所,穆恩斯特大学)

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL

Comments 9 pages, 4 figures, accepted at NCME AIME 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05473 2025-09-25 cs.MM cs.AI cs.SD eess.AS 79%

Embedding Alignment in Code Generation for Audio

Sam Kouteili, Hiren Madhu, George Typaldos, Mark Santolucito

机构 * Yale University(耶鲁大学) Columbia University(哥伦比亚大学)

专题命中 其他安全 :alignment(title,abstract);分类 cs.AI

Comments Accepted to NeurIPS 2025 AI4Music Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19472 2025-09-25 eess.SY cs.SY 71%

Automatic and Scalable Safety Verification using Interval Reachability with Subspace Sampling

Brendan Gould, Akash Harapanahalli, Samuel Coogan

专题命中 其他安全 :safety(title)

Comments 6 pages, 3 figures. Updated to correct a small error in the dynamics presented in equation (12)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.20208 2025-09-25 cs.CL cs.AI cs.DB 62%

Play by the Type Rules: Inferring Constraints for LLM Functions in Declarative Programs

Parker Glenn, Alfy Samuel, Daben Liu

机构 * Capital One

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.20051 2025-09-25 cs.LG cs.AI 62%

One Filters All: A Generalist Filter for State Estimation

Shiqi Liu, Wenhan Cao, Chang Liu, Zeyu He, Tianyi Zhang, Shengbo Eben Li

机构 * School of Vehicle and Mobility, Tsinghua University(清华大学车辆与移动系统学院) College of Engineering, Peking University(北京大学工程学院) College of AI, Tsinghua University(清华大学人工智能学院)

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏