arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-07-30 至 2025-07-30 共收录 42 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 偏好对齐 4 篇

2507.20133 2025-07-30 cs.CL cs.AI cs.LG 82%

Sem-DPO: Mitigating Semantic Inconsistency in Preference Optimization for Prompt Engineering

Anas Mohamed, Azal Ahmad Khan, Xinran Wang, Ahmad Faraz Khan, Shuwen Ge, Saman Bahzad Khan, Ayaan Ahmad, Ali Anwar

机构 * University of Minnesota(明尼苏达大学) Virginia Tech(弗吉尼亚理工大学) Xi’an University of Technology(西安理工大学) Lahore University of Management Sciences(拉合尔管理科学大学) University of California, Santa Cruz(加州大学圣克鲁兹分校)

专题命中 偏好对齐 :DPO(title,abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21931 2025-07-30 cs.CL cs.AI 62%

Post-Training Large Language Models via Reinforcement Learning from Self-Feedback

Carel van Niekerk, Renato Vukovic, Benjamin Matthias Ruppik, Hsien-chin Lin, Milica Gašić

机构 * Heinrich Heine Universität(海因里希·海因大学)

专题命中 偏好对齐 :RLHF(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.17107 2025-07-30 cs.LG cs.AI 62%

Reinforcement Learning Fine-Tunes a Sparse Subnetwork in Large Language Models

Andrii Balashov

专题命中 偏好对齐 :DPO(abstract);分类 cs.AI、cs.LG

Comments The manuscript has been withdrawn due to significant overlap in methodology and results with a prior work (arXiv:2505.11711) that we were not aware of at the time of submission. To maintain academic integrity and avoid redundancy in the literature, we have chosen to withdraw this version

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.18013 2025-07-30 cs.CL 57%

Technical Report of TeleChat2, TeleChat2.5 and T1

Zihan Wang, Xinzhang Liu, Yitong Yao, Chao Wang, Yu Zhao, Zhihao Yang, Wenmin Deng, Kaipeng Jia, Jiaxin Peng, Yuyao Huang, Sishi Xiong, Zhuo Jiang, Kaidong Yu, Xiaohui Hu, Fubei Yao, Ruiyu Fang, Zhuoru Jiang, Ruiting Song, Qiyi Xie, Rui Xue, Xuewei He, Yanlei Xue, Zhu Yuan, Zhaoxi Zhang, Zilu Huang, Shiquan Wang, Xin Wang, Hanming Wu, Mingyuan Wang, Xufeng Zhan, Yuhan Sun, Zhaohu Xing, Yuhao Jiang, Bingkai Yang, Shuangyong Song, Yongxiang Li, Zhongjiang He, Xuelong Li

专题命中 偏好对齐 :DPO(abstract);分类 cs.CL

Comments 32 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 安全训练 5 篇

2507.21132 2025-07-30 cs.AI cs.CY cs.LG 75%

Can You Trust an LLM with Your Life-Changing Decision? An Investigation into AI High-Stakes Responses

Joshua Adrian Cahyono, Saran Subramanian

专题命中 安全训练 :alignment(abstract);safety(abstract);分类 cs.AI、cs.CY、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21133 2025-07-30 cs.CR cs.AI 70%

Analysis of Threat-Based Manipulation in Large Language Models: A Dual Perspective on Vulnerabilities and Performance Enhancement Opportunities

Atil Samancioglu

机构 * Atil Samancioglu(独立研究者)

专题命中 安全训练 :safety(abstract);AI safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18172 2025-07-30 cs.CR cs.LG 57%

GenAI Security: Outsmarting the Bots with a Proactive Testing Framework

Sunil Kumar Jang Bahadur, Gopala Dhar, Lavi Nigam

机构 * AI \& GenAI Specialist Cloud GTM Google Mumbai, India AI Engineer, AI Services Google Cloud Consulting (GCC) Google Mumbai, India Industry Solutions Google Gurugram, India

专题命中 安全训练 :prompt injection(abstract);分类 cs.LG

Comments IEEE CAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21619 2025-07-30 cs.CV 50%

EMIT: Enhancing MLLMs for Industrial Anomaly Detection via Difficulty-Aware GRPO

Wei Guan, Jun Lan, Jian Cao, Hao Tan, Huijia Zhu, Weiqiang Wang

专题命中 安全训练 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21547 2025-07-30 math.OC cs.RO cs.SY eess.SY 50%

Decentralized Modeling of Vehicular Maneuvers and Interactions at Urban Junctions

Saeed Rahmani, Simeon C. Calvert, Bart van Arem

机构 * Delft University of Technology(代尔夫特理工大学)

专题命中 安全训练 :safety(abstract)

Comments Manuscript under review

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 越狱攻击 3 篇

2507.21820 2025-07-30 cs.CV 85%

Anyone Can Jailbreak: Prompt-Based Attacks on LLMs and T2Is

Ahmed B Mustafa, Zihan Ye, Yang Lu, Michael P Pound, Shreyank N Gowda

机构 * School of Computer Science, University of Nottingham(诺丁汉大学计算机科学学院) Department of Intelligent Science, Xi’an Jiaotong-Liverpool University(西安交通大学利物浦大学智能科学系) School of Informatics, Xiamen University(厦门大学信息学院)

专题命中 越狱攻击 :jailbreak(title,abstract);alignment(abstract);safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.22037 2025-07-30 cs.CR cs.AI 70%

Secure Tug-of-War (SecTOW): Iterative Defense-Attack Training with Reinforcement Learning for Multimodal Model Security

Muzhi Dai, Shixuan Liu, Zhiyuan Zhao, Junyu Gao, Hao Sun, Xuelong Li

机构 * Institute of Artificial Intelligence (TeleAI), China Telecom, China(人工智能研究院(TeleAI),中国电信,中国) Northwestern Polytechnical University(西北工业大学) China Telecom, China(中国电信,中国)

专题命中 越狱攻击 :safety(abstract);jailbreak(abstract);分类 cs.AI

Comments 10 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21182 2025-07-30 cs.CR cs.AI 70%

SDD: Self-Degraded Defense against Malicious Fine-tuning

Zixuan Chen, Weikai Lu, Xin Lin, Ziqian Zeng

机构 * Zixuan Chen(陈子轩) Weikai Lu(卢伟凯) Xin Lin(林鑫) Ziqian Zeng(曾子谦)

专题命中 越狱攻击 :alignment(abstract);safety(abstract);分类 cs.AI

Comments Accepted by ACL2025

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 红队测试 1 篇

2507.21061 2025-07-30 cs.CR cs.CY 84%

Security practices in AI development

Petr Spelda, Vit Stritecky

专题命中 红队测试 :alignment(abstract);safety(abstract);red teaming(abstract);trustworthy(abstract)

Comments 11 pages

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 幻觉与事实性 2 篇

2506.18985 2025-07-30 cs.CV cs.AI 61%

GLIMPSE: Holistic Cross-Modal Explainability for Large Vision-Language Models

Guanxi Shen

机构 * Georgia Institute of Technology(佐治亚理工学院)

专题命中 幻觉与事实性 :alignment(abstract,comments);分类 cs.AI

Comments Keywords: Explainable Computer Vision, Large Vision-Language Models, AI Interpretability, Explainable AI, Visual Saliency, Attribution Maps, Cross-Modal Attribution, Human Attention Alignment, AI Transparency

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21406 2025-07-30 cs.AI 57%

Shapley Uncertainty in Natural Language Generation

Meilin Zhu, Gaojie Jin, Xiaowei Huang, Lijun Zhang

机构 * University of Exeter(埃克塞特大学) University of Liverpool(利物浦大学) ISCAS(国际信息科学协会)

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏

6. 安全评测 16 篇

2501.13818 2025-07-30 cs.AI cs.CV cs.LG 87%

Ensuring Medical AI Safety: Interpretability-Driven Detection and Mitigation of Spurious Model Behavior and Associated Data

Frederik Pahde, Thomas Wiegand, Sebastian Lapuschkin, Wojciech Samek

机构 * Fraunhofer Heinrich Hertz Institut(弗劳恩霍夫 Heinrich Hertz 研究所) Technische Universität Berlin(柏林技术大学) Berlin Institute for the Foundations of Learning and Data (BIFOLD)(柏林学习与数据基础研究所(BIFOLD))

专题命中 安全评测 :safety(title,abstract);AI safety(title);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21169 2025-07-30 cs.CY cs.AI 84%

Trustworthy AI: UK Air Traffic Control Revisited

Rob Procter, Mark Rouncefield

专题命中 安全评测 :trustworthy(title,abstract);safety(abstract);分类 cs.AI、cs.CY

Comments 6 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21637 2025-07-30 cs.AI 83%

Self-Aware Safety Augmentation: Leveraging Internal Semantic Understanding to Enhance Safety in Vision-Language Models

Wanying Wang, Zeyu Ma, Han Zheng, Xin Tan, Mingang Chen

机构 * Shanghai Key Laboratory of Computer Software Testing and Evaluating(上海软件测试与评估 key laboratory) Shanghai Normal University(上海Normal University) TrustAI Pte. Ltd. East China Normal University(东华师范大学)

专题命中 安全评测 :safety(title,abstract);alignment(abstract);分类 cs.AI

Comments Accepted by ACM Multimedia 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21782 2025-07-30 cs.CL 79%

The Problem with Safety Classification is not just the Models

Sowmya Vajjala

专题命中 安全评测 :safety(title,abstract);分类 cs.CL

Comments Pre-print, Short paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21741 2025-07-30 cs.CV cs.MM 78%

MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces

Shaojun E, Yuchen Yang, Jiaheng Wu, Yan Zhang, Tiejun Zhao, Ziyan Chen

机构 * Global Tone Communication Technology Co., Ltd.(全球 tone 通信技术有限公司) Faculty of computing, Harbin Institute of Technology(哈尔滨工业大学计算机学院) School of Computer Science and Technology, Beijing Jiaotong University(北京交通大学计算机科学与技术学院)

专题命中 安全评测 :alignment(title,abstract)

Comments 9 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21145 2025-07-30 cs.CR 78%

Leveraging Trustworthy AI for Automotive Security in Multi-Domain Operations: Towards a Responsive Human-AI Multi-Domain Task Force for Cyber Social Security

Vita Santa Barletta, Danilo Caivano, Gabriel Cellammare, Samuele del Vescovo, Annita Larissa Sciacovelli

专题命中 安全评测 :trustworthy(title,abstract)

Comments 13 pages, 6 figures, 1 table

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.22034 2025-07-30 cs.AI cs.CL cs.LG 67%

UserBench: An Interactive Gym Environment for User-Centric Agents

Cheng Qian, Zuxin Liu, Akshara Prabhakar, Zhiwei Liu, Jianguo Zhang, Haolin Chen, Heng Ji, Weiran Yao, Shelby Heinecke, Silvio Savarese, Caiming Xiong, Huan Wang

机构 * Salesforce AI Research(Salesforce AI研究部) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments 25 Pages, 17 Figures, 6 Tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21815 2025-07-30 cs.CL cs.CY 62%

HRIPBench: Benchmarking LLMs in Harm Reduction Information Provision to Support People Who Use Drugs

Kaixuan Wang, Chenxin Diao, Jason T. Jacques, Zhongliang Guo, Shuai Zhao

机构 * University of St. Andrews(圣安德鲁大学) University of Edinburgh(爱丁堡大学) Nanyang Technological University(南洋理工大学)

专题命中 安全评测 :safety(abstract);分类 cs.CL、cs.CY

Comments 15 pages, 5 figures, 12 tables, a dataset

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21504 2025-07-30 cs.LG cs.AI 62%

Evaluation and Benchmarking of LLM Agents: A Survey

Mahmoud Mohammadi, Yipeng Li, Jane Lo, Wendy Yip

机构 * SAP Labs(SAP实验室)

专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21340 2025-07-30 cs.CL cs.AI cs.DB cs.IR 62%

StructText: A Synthetic Table-to-Text Approach for Benchmark Generation with Multi-Dimensional Evaluation

Satyananda Kashyap, Sola Shirai, Nandana Mihindukulasooriya, Horst Samulowitz

机构 * IBM Research(IBM研究院)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

Comments Data available: https://huggingface.co/datasets/ibm-research/struct-text and code available at: https://github.com/ibm/struct-text

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21188 2025-07-30 cs.LG cs.AI 62%

Embeddings to Diagnosis: Latent Fragility under Agentic Perturbations in Clinical LLMs

Raj Krishnan Vijayaraj

机构 * Independent Researcher(独立研究者)

专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.06123 2025-07-30 cs.CR cs.AI cs.CV cs.LG 62%

Adversarial attacks and defenses in explainable artificial intelligence: A survey

Hubert Baniecki, Przemyslaw Biecek

机构 * University of Warsaw(华沙大学) Warsaw University of Technology(华沙理工大学)

专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.LG

Comments Accepted by Information Fusion

Journal ref Information Fusion, vol. 107, 102303, 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21488 2025-07-30 cs.AI 57%

Learning to Imitate with Less: Efficient Individual Behavior Modeling in Chess

Zhenwei Tang, Difan Jiao, Eric Xue, Reid McIlroy-Young, Jon Kleinberg, Siddhartha Sen, Ashton Anderson

机构 * University of Toronto(多伦多大学) Harvard University(哈佛大学) Cornell University(康奈尔大学) Microsoft Research(微软研究院)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21076 2025-07-30 cs.CY 57%

Making a Case for Research Collaboration Between Artificial Intelligence and Operations Research Experts

Radhika Kulkarni, Gianluca Brero, Yu Ding, Swati Gupta, Sven Koenig, Ramayya Krishnan, Thiago Serra, Phebe Vayanos, Segev Wasserkrug, Holly Wiberg

专题命中 安全评测 :trustworthy(abstract);分类 cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.03274 2025-07-30 cs.AI 57%

A Scalable Approach to Probabilistic Neuro-Symbolic Robustness Verification

Vasileios Manginas, Nikolaos Manginas, Edward Stevinson, Sherwin Varghese, Nikos Katzouris, Georgios Paliouras, Alessio Lomuscio

机构 * Department of Computer Science and Leuven.AI KU Leuven Belgium(比利时列日大学计算机科学系) Department of Computing Imperial College London UK(伦敦帝国学院计算机系)

专题命中 安全评测 :safety(abstract);分类 cs.AI

Comments 19th Conference on Neurosymbolic Learning and Reasoning

详情

展开后加载摘要…

URL PDF HTML 收藏