arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-08-20 至 2025-08-20 共收录 22 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 偏好对齐 3 篇

2504.16438 2025-08-20 cs.LG cs.AI cs.CR cs.DC 62%

POPri: Private Federated Learning using Preference-Optimized Synthetic Data

Charlie Hou, Mei-Yu Wang, Yige Zhu, Daniel Lazar, Giulia Fanti

机构 * Pittsburgh Supercomputing Center, Pittsburgh, USA(匹兹堡超级计算中心) Department of ECE, Carnegie Mellon University, Pittsburgh, PA(电子工程系,卡内基梅隆大学)

专题命中 偏好对齐 :DPO(abstract);分类 cs.AI、cs.LG

Comments ICML 2025 camera-ready

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.13250 2025-08-20 cs.AI cs.CL cs.IR 62%

Explicit v.s. Implicit Memory: Exploring Multi-hop Complex Reasoning Over Personalized Information

Zeyu Zhang, Yang Zhang, Haoran Tan, Rui Li, Xu Chen

机构 * Gaoling School of Artificial Intelligence, Renmin University of China(中国人民大学人工智能学院) National University of Singapore(新加坡国立大学)

专题命中 偏好对齐 :alignment(abstract);分类 cs.CL、cs.AI

Comments 15 pages, 13 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.13189 2025-08-20 stat.ML cs.AI cs.LG 62%

Preference Models assume Proportional Hazards of Utilities

Chirag Nagpal

机构 * Meta Superintelligence Labs (MSL)(Meta超智能实验室)

专题命中 偏好对齐 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 安全训练 1 篇

2508.13608 2025-08-20 eess.SY cs.LG cs.SY math.OC 57%

Towards safe control parameter tuning in distributed multi-agent systems

Abdullah Tokmak, Thomas B. Schön, Dominik Baumann

专题命中 安全训练 :safety(abstract);分类 cs.LG

Comments Accepted to CDC 2025

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 越狱攻击 2 篇

2508.09288 2025-08-20 cs.CR cs.AI cs.CL 73%

Can AI Keep a Secret? Contextual Integrity Verification: A Provable Security Architecture for LLMs

Aayush Gupta

机构 * Aayush Gupta(独立研究者)

专题命中 越狱攻击 :jailbreak(abstract);prompt injection(abstract);分类 cs.CL、cs.AI

Comments 2 figures, 3 tables; code and certification harness: https://github.com/ayushgupta4897/Contextual-Integrity-Verification ; Elite-Attack dataset: https://huggingface.co/datasets/zyushg/elite-attack

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.02332 2025-08-20 cs.CR 50%

PII Jailbreaking in LLMs via Activation Steering Reveals Personal Information Leakage

Krishna Kanth Nakka, Xue Jiang, Dmitrii Usynin, Xuebing Zhou

专题命中 越狱攻击 :alignment(abstract)

Comments Preprint. V2 Updated with dataset filtering, benchmarking privacy evaluator and additional latent space visualizations

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 提示注入 1 篇

2508.13214 2025-08-20 cs.CR cs.AI 79%

Too Easily Fooled? Prompt Injection Breaks LLMs on Frustratingly Simple Multiple-Choice Questions

Xuyang Guo, Zekai Huang, Zhao Song, Jiahao Zhang

专题命中 提示注入 :prompt injection(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 幻觉与事实性 1 篇

2508.13564 2025-08-20 cs.CV cs.AI cs.LG cs.RO 62%

The 9th AI City Challenge

Zheng Tang, Shuo Wang, David C. Anastasiu, Ming-Ching Chang, Anuj Sharma, Quan Kong, Norimasa Kobori, Munkhjargal Gochoo, Ganzorig Batnasan, Munkh-Erdene Otgonbold, Fady Alnajjar, Jun-Wei Hsieh, Tomasz Kornuta, Xiaolong Li, Yilin Zhao, Han Zhang, Subhashree Radhakrishnan, Arihant Jain, Ratnesh Kumar, Vidya N. Murali, Yuxing Wang, Sameer Satish Pusegaonkar, Yizhou Wang, Sujit Biswas, Xunlei Wu, Zhedong Zheng, Pranamesh Chakraborty, Rama Chellappa

机构 * NVIDIA Corporation(NVIDIA公司) Santa Clara University(圣克拉拉大学) University at Albany, SUNY(纽约州立大学阿尔巴尼分校) Iowa State University(爱荷华州立大学) Woven by Toyota, Japan(日本丰田公司) United Arab Emirates University(阿拉伯联合酋长国大学) National Yang-Ming Chiao-Tung University(国家阳明交通大学) University of Macau(澳门大学) Indian Institute of Technology Kanpur(印度理工学院坎普尔分校) Johns Hopkins University(约翰霍普金斯大学)

专题命中 幻觉与事实性 :safety(abstract);分类 cs.AI、cs.LG

Comments Summary of the 9th AI City Challenge Workshop in conjunction with ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏

6. 安全评测 8 篇

2508.13179 2025-08-20 cs.CY cs.AI 88%

Toward an African Agenda for AI Safety

Samuel T. Segun, Rachel Adams, Ana Florido, Scott Timcke, Jonathan Shock, Leah Junck, Fola Adeleke, Nicolas Grossman, Ayantola Alayande, Jerry John Kponyo, Matthew Smith, Dickson Marfo Fosu, Prince Dawson Tetteh, Juliet Arthur, Stephanie Kasaon, Odilile Ayodele, Laetitia Badolo, Paul Plantinga, Michael Gastrow, Sumaya Nur Adan, Joanna Wiaterek, Cecil Abungu, Kojo Apeagyei, Luise Eder, Tegawende Bissyande

专题命中 安全评测 :safety(title,abstract);AI safety(title,abstract);分类 cs.AI、cs.CY

Comments 28 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.13787 2025-08-20 cs.MA cs.AI cs.NI 79%

BetaWeb: Towards a Blockchain-enabled Trustworthy Agentic Web

Zihan Guo, Yuanjian Zhou, Chenyi Wang, Linlin You, Minjie Bian, Weinan Zhang

机构 * Shanghai Innovation Institute(上海创新研究院) Sun Yat-sen University(中山大学) Zhejiang University(浙江大学) Shanghai Data Group Co., Ltd(上海数据集团有限公司) Shanghai Jiao Tong University(上海交通大学)

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI

Comments A technical report with 21 pages, 3 figures, and 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.13180 2025-08-20 cs.AI cs.LG 62%

Search-Time Data Contamination

Ziwen Han, Meher Mankikar, Julian Michael, Zifan Wang

机构 * Scale AI

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.13465 2025-08-20 cs.AI 57%

LM Agents May Fail to Act on Their Own Risk Knowledge

Yuzhi Tang, Tianxiao Li, Elizabeth Li, Chris J. Maddison, Honghua Dong, Yangjun Ruan

专题命中 安全评测 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.03666 2025-08-20 cs.RO cs.LG 57%

Hybrid Machine Learning Model with a Constrained Action Space for Trajectory Prediction

Alexander Fertig, Lakshman Balasubramanian, Michael Botsch

机构 * Technische Hochschule Ingolstadt, AImotion Bavaria(图腾工业大学,拜耳巴伐利亚人工智能公司) Technische Hochschule Ingolstadt, Research Center CARISSMA(图腾工业大学,CARISSMA研究中心)

专题命中 安全评测 :safety(abstract);分类 cs.LG

Comments Copyright 2025 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works

Journal ref 2025 IEEE Intelligent Vehicles Symposium (IV)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.13880 2025-08-20 cs.CV 50%

In-hoc Concept Representations to Regularise Deep Learning in Medical Imaging

Valentina Corbetta, Floris Six Dijkstra, Regina Beets-Tan, Hoel Kervadec, Kristoffer Wickstrøm, Wilson Silva

机构 * The Netherlands Cancer Institute(荷兰癌症研究所) Utrecht University(乌得勒支大学) Maastricht University(马斯特里赫特大学) University of Amsterdam(阿姆斯特丹大学) Amsterdam UMC(阿姆斯特丹大学医学中心) UiT The Arctic University of Norway(挪威北莫斯堡大学)

专题命中 安全评测 :trustworthy(abstract)

Comments 13 pages, 13 figures, 2 tables, accepted at PHAROS-AFE-AIMI Workshop in conjunction with the International Conference on Computer Vision (ICCV), 2025. This is the submitted manuscript with added link to the github repo, funding acknowledgments and author names and affiliations, and a correction to numbers in Table 1. Final version not published yet

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.13413 2025-08-20 cs.HC cs.SE 50%

Large Language Models as Visualization Agents for Immersive Binary Reverse Engineering

Dennis Brown, Samuel Mulder

专题命中 安全评测 :alignment(abstract)

Comments Accepted to IEEE VISSOFT 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06189 2025-08-20 cs.CV 50%

MA-CBP: A Criminal Behavior Prediction Framework Based on Multi-Agent Asynchronous Collaboration

Cheng Liu, Daou Zhang, Tingxu Liu, Yuhan Wang, Jinyang Chen, Yuexuan Li, Xinying Xiao, Chenbo Xin, Ziru Wang, Weichao Wu

专题命中 安全评测 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

7. AI治理与伦理 2 篇

2508.13960 2025-08-20 cs.GT cs.AI 70%

A Mechanism for Mutual Fairness in Cooperative Games with Replicable Resources -- Extended Version

Björn Filter, Ralf Möller, Özgür Lütfü Özçep

机构 * Institute for Humanities-Centered AI (CHAI), University of Hamburg, Germany(人文中心人工智能研究所(CHAI),汉堡大学)

专题命中 AI治理与伦理 :alignment(abstract);safety(abstract);分类 cs.AI

Comments This paper is the extended version of a paper accepted at the European Conference on Artificial Intelligence 2025 (ECAI 2025), providing the proof of the main theorem in the appendix

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.13743 2025-08-20 cs.CL 57%

Sycophancy under Pressure: Evaluating and Mitigating Sycophantic Bias via Adversarial Dialogues in Scientific QA

Kaiwei Zhang, Qi Jia, Zijian Chen, Wei Sun, Xiangyang Zhu, Chunyi Li, Dandan Zhu, Guangtao Zhai

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏

8. 其他安全 4 篇

2503.08549 2025-08-20 cs.AI cs.CL 62%

GoAI: Enhancing AI Students' Learning Paths and Idea Generation via Graph of AI Ideas

Xian Gao, Zongyun Zhang, Ting Liu, Yuzhuo Fu

机构 * Shanghai Jiao Tong University(上海交通大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

Comments Work in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.06723 2025-08-20 cs.HC cs.AI 57%

Script-Strategy Aligned Generation: Aligning LLMs with Expert-Crafted Dialogue Scripts and Therapeutic Strategies for Psychotherapy

Xin Sun, Jan de Wit, Zhuying Li, Jiahuan Pei, Abdallah El Ali, Jos A. Bosch

机构 * University of Amsterdam(阿姆斯特丹大学) National Institute of Informatics (NII)(日本信息处理研究所) Tilburg University(蒂尔堡大学) Southeast University(东南大学) Vrije Universiteit Amsterdam(阿姆斯特丹自由大学) Utrecht University(乌得勒支大学)

专题命中 其他安全 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.13901 2025-08-20 cs.RO cs.CV 50%

Multimodal Data Storage and Retrieval for Embodied AI: A Survey

Yihao Lu, Hao Tang

机构 * School of Economics and Management, South China Normal University(经济管理学院,华南师范大学) State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University(多媒体信息处理国家重点实验室,计算机学院,北京大学)

专题命中 其他安全 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.14658 2025-08-20 cs.CV 50%

EmoSEM: Segment and Explain Emotion Stimuli in Visual Art

Jing Zhang, Dan Guo, Zhangbin Li, Meng Wang

机构 * Hefei University of Technology(合肥工业大学)

专题命中 其他安全 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏