arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-10-14 至 2025-10-14 共收录 97 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 36 篇

2510.10085 2025-10-14 cs.CR cs.AI cs.LG 88%

Pharmacist: Safety Alignment Data Curation for Large Language Models against Harmful Fine-tuning

Guozhi Liu, Qi Mu, Tiansheng Huang, Xinhua Wang, Li Shen, Weiwei Lin, Zhang Li

机构 * School of Computer Science and Engineering, South China University of Technology(华南理工大学计算机科学与工程学院) School of Computer Science at Georgia Institute of Technology(佐治亚理工学院计算机科学系) School of Cyber Science and Technology, Shenzhen Campus of Sun Yat-sen University(中山大学深圳校区网络科学与技术学院) Second Affiliated Hospital of Guangzhou University of Chinese Medicine(广州中医药大学第二附属医院) China and Pengcheng Laboratory(中 Pengcheng 实验室)

专题命中 安全评测 :alignment(title,abstract);safety(title,abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.24394 2025-10-14 cs.CY cs.AI 87%

The 2025 OpenAI Preparedness Framework does not guarantee any AI risk mitigation practices: a proof-of-concept for affordance analyses of AI safety policies

Sam Coggins, Alexander K. Saeri, Katherine A. Daniell, Lorenn P. Ruster, Jessie Liu, Jenny L. Davis

专题命中 安全评测 :safety(title,abstract);AI safety(title);分类 cs.AI、cs.CY

Comments 19 pages, 5 tables, 1 figure; minor ambiguities clarified, typos corrected, author affiliations added

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.12661 2025-10-14 cs.LG cs.CL cs.CV 87%

VLMGuard-R1: Proactive Safety Alignment for VLMs via Reasoning-Driven Prompt Optimization

Menglan Chen, Xianghe Pang, Jingjing Dong, WenHao Wang, Yaxin Du, Siheng Chen

机构 * Shanghai Jiao Tong University(上海交通大学)

专题命中 安全评测 :safety(title,abstract);alignment(title);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11314 2025-10-14 cs.CL 79%

Template-Based Text-to-Image Alignment for Language Accessibility: A Study on Visualizing Text Simplifications

Belkiss Souayed, Sarah Ebling, Yingqiang Gao

机构 * University of Zurich(苏黎世大学)

专题命中 安全评测 :alignment(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10677 2025-10-14 cs.CL 79%

Unlocking LLM Safeguards for Low-Resource Languages via Reasoning and Alignment with Minimal Training Data

Zhuowei Chen, Bowei Zhang, Nankai Lin, Tian Hou, Lianxi Wang

机构 * Guangdong University of Foreign Studies(广东外语外贸大学) Guangzhou Key Laboratory of Multilingual Intelligent Processing(广州多语智能处理重点实验室) University of Pittsburgh(匹兹堡大学)

专题命中 安全评测 :alignment(title,abstract);分类 cs.CL

Comments Accepted to MRL Workshop at EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10086 2025-10-14 cs.RO 78%

Beyond ADE and FDE: A Comprehensive Evaluation Framework for Safety-Critical Prediction in Multi-Agent Autonomous Driving Scenarios

Feifei Liu, Haozhe Wang, Zejun Wei, Qirong Lu, Yiyang Wen, Xiaoyu Tang, Jingyan Jiang, Zhijian He

机构 * South China Normal University(华南师范大学) Shenzhen Technology University(深圳技术大学)

专题命中 安全评测 :safety(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.17040 2025-10-14 cs.CV 78%

Multimodal Alignment and Fusion: A Survey

Songtao Li, Hao Tang

机构 * Peking University(北京大学) Northeastern University(东北大学) Sydney Smart Technology College(悉尼智能技术学院) School of Computer Science, Peking University(北京大学计算机学院) The State Key Laboratory of Multimedia Information Processing(多媒体信息处理国家重点实验室)

专题命中 安全评测 :alignment(title,abstract)

Comments Accepted to IJCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10426 2025-10-14 cs.CV cs.AI 70%

Taming a Retrieval Framework to Read Images in Humanlike Manner for Augmenting Generation of MLLMs

Suyang Xi, Chenxi Yang, Hong Ding, Yiqing Ni, Catherine C. Liu, Yunhao Liu, Chengqi Zhang

机构 * Emory University(埃默里大学) University of Electronic Science and Technology of China(电子科技大学) University of Illinois Chicago(伊利诺伊大学香槟分校) The Hong Kong Polytechnic University(香港理工大学)

专题命中 安全评测 :alignment(abstract);trustworthy(abstract);分类 cs.AI

Comments 12 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.23042 2025-10-14 cs.CV cs.AI cs.LG cs.MM cs.RO 62%

Goal-Based Vision-Language Driving

Santosh Patapati, Trisanth Srinivasan

机构 * Dept. of HCI(人机交互系) Cyrion Labs(Cyrion实验室)

专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.LG

Comments 6 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.19981 2025-10-14 cs.LG cs.CL 62%

Accurate and Diverse LLM Mathematical Reasoning via Automated PRM-Guided GFlowNets

Adam Younsi, Ahmed Attia, Abdalgader Abubaker, Mohamed El Amine Seddik, Hakim Hacid, Salem Lahlou

机构 * Technology Innovation Institute(技术创新研究所) Mohamed Bin Zayed University of Artificial Intelligence(莫罕默德·本·扎耶德人工智能大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.09479 2025-10-14 cs.AI cs.CL 62%

Draw with Thought: Unleashing Multimodal Reasoning for Scientific Diagram Generation

Zhiqing Cui, Jiahao Yuan, Hanqing Wang, Yanshu Li, Chenxu Du, Zhenglong Ding

机构 * Nanjing University of Information Science \& Technology Nanjing China East China Normal University Shanghai China The Hong Kong University of Science Brown University Providence America Southwest Jiaotong University Chengdu China Nanjing University of Information Science \& Technology East China Normal University Brown University Southwest Jiaotong University

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

Comments 10 pages, 5 figures, accepted to appear in the Proceedings of the 33rd ACM International Conference on Multimedia (MM '25)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.09738 2025-10-14 cs.CL cs.AI 62%

Judge's Verdict: A Comprehensive Analysis of LLM Judge Capability Through Human Agreement

Steve Han, Gilberto Titericz Junior, Tom Balough, Wenfei Zhou

机构 * NVIDIA Corporation(NVIDIA公司)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

Comments 10 pages, 1 figure, 4 tables, under review as a conference paper at ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.10983 2025-10-14 cs.LG cs.AI cs.CR 62%

GenoArmory: A Unified Evaluation Framework for Adversarial Attacks on Genomic Foundation Models

Haozheng Luo, Chenghao Qiu, Yimin Wang, Shang Wu, Jiahao Yu, Zhenyu Pan, Weian Mao, Haoyang Fang, Hao Xu, Han Liu, Binghui Wang, Yan Chen

机构 * Department of Computer Science, Northwestern University(西北大学计算机科学系) Department of Computer Science and Engineering, Texas A&M University(德克萨斯农工大学计算机科学与工程系) Department of Computer Science, Illinois Institute of Technology(伊利诺伊理工学院计算机科学系) Department of Statistics and Data Science, Northwestern University(西北大学统计与数据科学系) Department of Computer Science and Engineering, University of Michigan(密歇根大学计算机科学与工程系) Department of Electrical Engineering and Computer Science, Massachusetts Institute of Technology(麻省理工学院电气工程与计算机科学系) Department of Computer Science, Carnegie Mellon University(卡内基梅隆大学计算机科学系) Department of Medicine at Brigham and Women’s Hospital and Harvard Medical School(布里格姆和妇女医院及哈佛医学院医学系)

专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.14858 2025-10-14 cs.AI cs.CL 62%

Retrieval is Not Enough: Enhancing RAG Reasoning through Test-Time Critique and Optimization

Jiaqi Wei, Hao Zhou, Xiang Zhang, Di Zhang, Zijie Qiu, Wei Wei, Jinzhe Li, Wanli Ouyang, Siqi Sun

机构 * Zhejiang University(浙江大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) South China University of Technology(华南理工大学) University of British Columbia(不列颠哥伦比亚大学) Fudan University(复旦大学) University of Hong Kong(香港大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

Comments Accepted by NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11688 2025-10-14 cs.CR cs.AI 57%

PACEbench: A Framework for Evaluating Practical AI Cyber-Exploitation Capabilities

Zicheng Liu, Lige Huang, Jie Zhang, Dongrui Liu, Yuan Tian, Jing Shao

机构 * Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Shanghai Jiao Tong University(上海交通大学) Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

Comments Project webpage available at https://pacebench.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11389 2025-10-14 cs.CL 57%

Beyond Survival: Evaluating LLMs in Social Deduction Games with Human-Aligned Strategies

Zirui Song, Yuan Huang, Junchang Liu, Haozhe Luo, Chenxi Wang, Lang Gao, Zixiang Xu, Mingfei Han, Xiaojun Chang, Xiuying Chen

机构 * Mohamed bin Zayed University of Artificial Intelligence (MBZUAI)(穆罕默德·本·扎耶德人工智能大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

Comments 34 pages, 32figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11313 2025-10-14 cs.AI 57%

Automated Skill Decomposition Meets Expert Ontologies: Bridging the Granularity Gap with LLMs

Le Ngoc Luyen, Marie-Hélène Abel

机构 * Université de technologie de Compiègne, CNRS, Heudiasyc(法国图卢兹技术大学、CNRS、Heudiasyc)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11202 2025-10-14 cs.LG cs.CR 57%

Evaluating Line-level Localization Ability of Learning-based Code Vulnerability Detection Models

Marco Pintore, Giorgio Piras, Angelo Sotgiu, Maura Pintor, Battista Biggio

机构 * University of Cagliari, Italy(卡利亚里大学) CINI, Italy(CINI)

专题命中 安全评测 :alignment(abstract);分类 cs.LG

Comments Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11151 2025-10-14 cs.CL cs.CR 57%

TypePilot: Leveraging the Scala Type System for Secure LLM-generated Code

Alexander Sternfeld, Andrei Kucharavy, Ljiljana Dolamic

机构 * Institute of Entrepreneurship & Management, HES-SO Le Foyer, Techno-Pôle 1 Sierre, Switzerland(创业与管理学院,HES-SO莱福院,技术园区1,瑞士) Institute of Informatics, HES-SO Techno-Pôle 3 Sierre, Switzerland(信息学院,HES-SO技术园区3,瑞士) Cyber-Defence Campus armasuisse, Science and Technology Thun, Switzerland(网络安全校区,armasuisse,科技,瑞士)

专题命中 安全评测 :safety(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11131 2025-10-14 cs.SI cs.CY 57%

SocioBench: Modeling Human Behavior in Sociological Surveys with Large Language Models

Jia Wang, Ziyu Zhao, Tingjuntao Ni, Zhongyu Wei

专题命中 安全评测 :alignment(abstract);分类 cs.CY

Comments Accepted by EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10719 2025-10-14 cs.SD cs.AI 57%

SS-DPPN: A self-supervised dual-path foundation model for the generalizable cardiac audio representation

Ummy Maria Muna, Md Mehedi Hasan Shawon, Md Jobayer, Sumaiya Akter, Md Rakibul Hasan, Md. Golam Rabiul Alam

机构 * Department of Computer Science and Engineering(计算机科学与工程系) BRAC University(布拉克大学) Department of Electricial and Electronic Engineering(电气与电子工程系) Department of Biomedical Engineering(生物医学工程系) Linköping University(林肯堡大学) Department of Electrical and Computer Engineering(电气与计算机工程系) University of Maryland(马里兰大学) School of Electrical Engineering, Computing and Mathematical Sciences(电气工程、计算与数学科学学院)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.09116 2025-10-14 cs.CL 57%

DITING: A Multi-Agent Evaluation Framework for Benchmarking Web Novel Translation

Enze Zhang, Jiaying Wang, Mengxi Xiao, Jifei Liu, Ziyan Kuang, Rui Dong, Eric Dong, Sophia Ananiadou, Min Peng, Qianqian Xie

机构 * School of Artificial Intelligence, Wuhan University(武汉大学人工智能学院) Center for Language and Information Research, Wuhan University(武汉大学语言信息研究中心) Jiangxi Normal University(江西师范大学) The University of Manchester(曼彻斯特大学) Yunnan Trrans Technology Co., Ltd.(云南翻译技术有限公司) Malvern College Chengdu(成都马尔伯学院)

专题命中 安全评测 :safety(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04852 2025-10-14 cs.SE cs.AI 57%

FreshBrew: A Benchmark for Evaluating AI Agents on Java Code Migration

Victor May, Diganta Misra, Yanqi Luo, Anjali Sridhar, Justine Gehring, Silvio Soares Ribeiro Junior

机构 * Google(谷歌) Max Planck Institut für Intelligente Systeme (MPI-IS)(马克斯·普朗克智能系统研究所) ELLIS Institute, Tübingen(图宾根ELLIS研究所) Salesforce Gologic Inc(Gologic公司)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

Comments 18 pages, 12 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.19073 2025-10-14 cs.CL 57%

MFTCXplain: A Multilingual Benchmark Dataset for Evaluating the Moral Reasoning of LLMs through Multi-hop Hate Speech Explanation

Jackson Trager, Francielle Vargas, Diego Alves, Matteo Guida, Mikel K. Ngueajio, Ameeta Agrawal, Yalda Daryani, Farzan Karimi-Malekabadi, Flor Miriam Plaza-del-Arco

机构 * University of Southern California(南加州大学) São Paulo State University(圣保罗州立大学) Saarland University(萨尔兰大学) University of Melbourne(墨尔本大学) Howard University(霍华德大学) Portland State University(波特兰州立大学) Leiden University(莱顿大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

Comments Jackson Trager and Francielle Vargas contributed equally

Journal ref Findings of the Association for Computational Linguistics: EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10224 2025-10-14 cs.CL cs.IR 57%

Text2Token: Unsupervised Text Representation Learning with Token Target Prediction

Ruize An, Richong Zhang, Zhijie Nie, Zhanyu Wu, Yanzhao Zhang, Dingkun Long

机构 * CCSE, School of Computer Science and Engineering, Beihang University, Beijing, China(计算机科学与工程学院,北京航空航天大学,北京,中国) Zhongguancun Laboratory, Beijing, China(中关村实验室,北京,中国) Shen Yuan Honors College, Beihang University, Beijing, China(神元荣誉学院,北京航空航天大学,北京,中国)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10199 2025-10-14 cs.HC cs.AI 57%

Revisiting Trust in the Era of Generative AI: Factorial Structure and Latent Profiles

Haocan Sun, Weizi Liu, Di Wu, Guoming Yu, Mike Yao

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.09620 2025-10-14 cs.CR cs.AI 57%

Toward a Unified Security Framework for AI Agents: Trust, Risk, and Liability

Jiayun Mo, Xin Kang, Tieyan Li, Zhongding Lei

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.06343 2025-10-14 cs.SE cs.AI cs.CR 57%

Leveraging Large Language Models for Cybersecurity Risk Assessment -- A Case from Forestry Cyber-Physical Systems

Fikret Mert Gultekin, Oscar Lilja, Ranim Khojah, Rebekka Wohlrab, Marvin Damschen, Mazen Mohamad

机构 * Chalmers University of Technology(楚德斯技术大学) University of Gothenburg(哥德堡大学) Carnegie Mellon University(卡内基梅隆大学) RISE Research Institutes of Sweden(瑞典RISE研究机构)

专题命中 安全评测 :safety(abstract);分类 cs.AI

Comments Accepted at Autonomous Agents in Software Engineering (AgenticSE) Workshop, co-located with ASE 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03418 2025-10-14 cs.AI cs.MA 57%

LegalWiz: A Multi-Agent Generation Framework for Contradiction Detection in Legal Documents

Ananya Mantravadi, Shivali Dalmia, Olga Pospelova, Abhishek Mukherji, Nand Dave, Anudha Mittal

机构 * Centific Amazon(亚马逊)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.14670 2025-10-14 cs.HC cs.AI 57%

StreetLens: Enabling Human-Centered AI Agents for Neighborhood Assessment from Street View Imagery

Jina Kim, Leeje Jang, Yao-Yi Chiang, Guanyu Wang, Michelle C. Pasco

机构 * University of Minnesota(明尼苏达大学)

专题命中 安全评测 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏