arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-08-22 至 2025-08-22 共收录 41 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 偏好对齐 6 篇

2412.19792 2025-08-22 cs.LG cs.CL cs.IT math.IT 84%

InfAlign: Inference-aware language model alignment

Ananth Balashankar, Ziteng Sun, Jonathan Berant, Jacob Eisenstein, Michael Collins, Adrian Hutter, Jong Lee, Chirag Nagpal, Flavien Prost, Aradhana Sinha, Ananda Theertha Suresh, Ahmad Beirami

机构 * Google DeepMind(谷歌DeepMind) Google Research(谷歌研究)

专题命中 偏好对齐 :alignment(title,abstract);RLHF(abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15388 2025-08-22 cs.IR 78%

TrackRec: Iterative Alternating Feedback with Chain-of-Thought via Preference Alignment for Recommendation

Yu Xia, Rui Zhong, Zeyu Song, Wei Yang, Junchen Wan, Qingpeng Cai, Chi Lu, Peng Jiang

专题命中 偏好对齐 :alignment(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15449 2025-08-22 cs.LG cs.AI 62%

Reliable Unlearning Harmful Information in LLMs with Metamorphosis Representation Projection

Chengcan Wu, Zeming Wei, Huanran Chen, Yinpeng Dong, Meng Sun

机构 * Peking University(北京大学) Tsinghua University(清华大学)

专题命中 偏好对齐 :safety(abstract);分类 cs.AI、cs.LG

Comments 10 pages, 9 figures, Under review as a full paper at AAAI 2026. A preliminary version is under review at the NeurIPS 2025 Workshop on Reliable ML from Unreliable Data

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15548 2025-08-22 cs.AI 57%

DeepThink3D: Enhancing Large Language Models with Programmatic Reasoning in Complex 3D Situated Reasoning Tasks

Jiayi Song, Rui Wan, Lipeng Ma, Weidong Yang, Qingyuan Zhou, Yixuan Li, Ben Fei

专题命中 偏好对齐 :DPO(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.14951 2025-08-22 cs.CL 57%

Improving LLMs for Machine Translation Using Synthetic Preference Data

Dario Vajda, Domen Vreš, Marko Robnik-Šikonja

机构 * University of Ljubljana, Faculty of Computer and Information Science(卢布尔雅那大学计算机与信息科学系)

专题命中 偏好对齐 :DPO(abstract);分类 cs.CL

Comments Paper with individual presentation at LUHME workshop at ECAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15146 2025-08-22 cs.HC 50%

QueryGenie: Making LLM-Based Database Querying Transparent and Controllable

Longfei Chen, Shenghan Gao, Shiwei Wang, Ken Lin, Yun Wang, Quan Li

专题命中 偏好对齐 :alignment(abstract)

Comments Accepted by The 38th Annual ACM Symposium on User Interface Software and Technology (UIST Adjunct '25), September 28-October 1, 2025, Busan, Republic of Korea

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 安全训练 1 篇

2508.15068 2025-08-22 cs.AI 70%

S3LoRA: Safe Spectral Sharpness-Guided Pruning in Adaptation of Agent Planner

Shuang Ao, Gopal Rumchurn

机构 * Shuang Ao Gopal Rumchurn

专题命中 安全训练 :alignment(abstract);safety(abstract);分类 cs.AI

Comments 9 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 越狱攻击 4 篇

2508.15182 2025-08-22 cs.LG 85%

SafeLLM: Unlearning Harmful Outputs from Large Language Models against Jailbreak Attacks

Xiangman Li, Xiaodong Wu, Qi Li, Jianbing Ni, Rongxing Lu

机构 * Department of Electrical and Computer Engineering, Queen’s University(电气与计算机工程系,皇后大学) School of Computing, Queen’s University(计算机学院,皇后大学)

专题命中 越狱攻击 :jailbreak(title,abstract);alignment(abstract);safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15283 2025-08-22 cs.IR cs.CL 57%

Adversarial Attacks against Neural Ranking Models via In-Context Learning

Amin Bigdeli, Negar Arabzadeh, Ebrahim Bagheri, Charles L. A. Clarke

机构 * University of Waterloo(多伦多大学) University of California, Berkeley(加州大学伯克利分校) University of Toronto(多伦多大学)

专题命中 越狱攻击 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.19521 2025-08-22 cs.CR 50%

Security Steerability is All You Need

Itay Hazan, Idan Habler, Ron Bitton, Itsik Mantin

专题命中 越狱攻击 :prompt injection(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.06154 2025-08-22 cs.CV 50%

GLOV: Guided Large Language Models as Implicit Optimizers for Vision Language Models

M. Jehanzeb Mirza, Mengjie Zhao, Zhuoyuan Mao, Sivan Doveh, Wei Lin, Paul Gavrikov, Michael Dorkenwald, Shiqi Yang, Saurav Jha, Hiromi Wakaki, Yuki Mitsufuji, Horst Possegger, Rogerio Feris, Leonid Karlinsky, James Glass

机构 * MIT CSAIL(麻省理工学院计算机科学与人工智能实验室) Sony(索尼) Weizmann Institute of Science(魏茨曼科学研究院) JKU(约翰纳斯堡大学) Tübingen AI Center(图宾根人工智能中心) UVA(乌得勒支大学) UNSW(新南威尔士大学) TU Graz(格拉茨技术大学) MIT-IBM(麻省理工学院-IBM)

专题命中 越狱攻击 :safety(abstract)

Comments Code: https://github.com/jmiemirza/GLOV

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 红队测试 1 篇

2508.12504 2025-08-22 cs.HC 78%

Organization Matters: A Qualitative Study of Organizational Dynamics in Red Teaming Practices for Generative AI

Bixuan Ren, EunJeong Cheon, Jianghui Li

专题命中 红队测试 :red teaming(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 提示注入 1 篇

2508.15310 2025-08-22 cs.CR cs.AI cs.CL 81%

IPIGuard: A Novel Tool Dependency Graph-Based Defense Against Indirect Prompt Injection in LLM Agents

Hengyu An, Jinghuai Zhang, Tianyu Du, Chunyi Zhou, Qingming Li, Tao Lin, Shouling Ji

专题命中 提示注入 :prompt injection(title,abstract);分类 cs.CL、cs.AI

Comments EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏

6. 隐私与版权 1 篇

2306.01334 2025-08-22 cs.LG cs.AI 62%

Federated Domain Generalization: A Survey

Ying Li, Xingwei Wang, Rongfei Zeng, Praveen Kumar Donta, Ilir Murturi, Min Huang, Schahram Dustdar

机构 * College of Computer Science and Engineering, Northeastern University(计算机科学与工程学院,东北大学) Distributed Systems Group, TU Wien(分布式系统组,TU Wien) College of Software, Northeastern University(软件学院,东北大学) College of Information Science and Engineering, Northeastern University(信息科学与工程学院,东北大学)

专题命中 隐私与版权 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏

7. 安全评测 20 篇

2508.15192 2025-08-22 cs.AI cs.CL 81%

LLM4Sweat: A Trustworthy Large Language Model for Hyperhidrosis Support

Wenjie Lin, Jin Wei-Kocsis

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.22146 2025-08-22 cs.CV cs.AI cs.CL q-bio.NC 81%

Flexible Tool Selection through Low-dimensional Attribute Alignment of Vision and Language

Guangfu Hao, Haojie Wen, Liangxuan Guo, Yang Chen, Yanchao Bi, Shan Yu

机构 * Laboratory of Brain Atlas and Brain-inspired Intelligence, Institute of Automation Chinese Academy of Sciences (CASIA)(中国科学院自动化研究所脑图谱与类脑智能实验室) School of Artificial Intelligence, University of Chinese Academy of Sciences (UCAS)(中国科学院大学人工智能学院) School of Systems Science, Beijing Normal University(北京师范大学系统科学学院) School of Psychological and Cognitive Sciences & Beijing Key Laboratory of Behavior and Mental Health, Peking University(北京大学心理与认知科学学院) IDG/McGovern Institute for Brain Research, Peking University(北京大学IDG/ McGovern脑科学研究院) Institute for Artificial Intelligence & Key Laboratory of Machine Perception (Ministry of Education), Peking University(北京大学人工智能研究所) School of Future Technology, University of Chinese Academy of Sciences (UCAS)(中国科学院大学未来技术学院)

专题命中 安全评测 :alignment(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15588 2025-08-22 cs.AI 79%

A Dynamical Systems Framework for Reinforcement Learning Safety and Robustness Verification

Ahmed Nasir, Abdelhafid Zenati

机构 * Engineering Department, School of Science and Technology (SST), City University of London(伦敦城市大学科学与技术学院工程系)

专题命中 安全评测 :safety(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15526 2025-08-22 cs.CL 79%

SafetyFlow: An Agent-Flow System for Automated LLM Safety Benchmarking

Xiangyang Zhu, Yuan Tian, Chunyi Li, Kaiwei Zhang, Wei Sun, Guangtao Zhai

机构 * Shanghai AI Lab(上海人工智能实验室) East China Normal University(东华大学)

专题命中 安全评测 :safety(title,abstract);分类 cs.CL

Comments Code and dataset are available at https://github.com/yangyangyang127/SafetyFlow

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15370 2025-08-22 cs.CL cs.AI 73%

Unveiling Trust in Multimodal Large Language Models: Evaluation, Analysis, and Mitigation

Yichi Zhang, Yao Huang, Yifan Wang, Yitong Sun, Chang Liu, Zhe Zhao, Zhengwei Fang, Huanran Chen, Xiao Yang, Xingxing Wei, Hang Su, Yinpeng Dong, Jun Zhu

机构 * Department of Computer Science and Technology, College of AI, Institute for AI, Tsinghua-Bosch Joint ML Center, THBI Lab, BNRist Center, Tsinghua University(计算机科学与技术系、人工智能学院、人工智能研究所、清华-博世联合机器学习中心、THBI实验室、BNRist中心、清华大学) Institute of Artificial Intelligence, Beihang University(人工智能研究院、北航) RealAI

专题命中 安全评测 :alignment(abstract);safety(abstract);分类 cs.CL、cs.AI

Comments For Appendix, please refer to arXiv:2406.07057

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.14925 2025-08-22 cs.CR cs.LG 70%

MCPTox: A Benchmark for Tool Poisoning Attack on Real-World MCP Servers

Zhiqiang Wang, Yichao Gao, Yanting Wang, Suyuan Liu, Haifeng Sun, Haoran Cheng, Guanquan Shi, Haohua Du, Xiangyang Li

专题命中 安全评测 :alignment(abstract);safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15418 2025-08-22 cs.CL cs.AI cs.LG cs.MM cs.SD 67%

LLaSO: A Foundational Framework for Reproducible Research in Large Language and Speech Model

Yirong Sun, Yizhong Geng, Peidong Wei, Yanjun Chen, Jinghan Yang, Rongfei Chen, Wei Zhang, Xiaoyu Shen

机构 * Ningbo Key Laboratory of Spatial Intelligence and Digital Derivative, Institute of Digital Twin, EIT(宁波空间智能与数字衍生关键实验室,数字孪生研究院,EIT) Logic Intelligence Technology(逻辑智能技术) BUPT(北京邮电大学) Xiamen University(厦门大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15085 2025-08-22 cs.CL cs.AI cs.IR cs.LG 67%

LongRecall: A Structured Approach for Robust Recall Evaluation in Long-Form Text

MohamamdJavad Ardestani, Ehsan Kamalloo, Davood Rafiei

机构 * Department of Computing Science University of Alberta(计算科学系阿尔伯塔大学) ServiceNow Research(ServiceNow研究)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15464 2025-08-22 cs.CL cs.AI 62%

RadReason: Radiology Report Evaluation Metric with Reasons and Sub-Scores

Yingshu Li, Yunyi Liu, Lingqiao Liu, Lei Wang, Luping Zhou

机构 * School of Electrical and Computer Engineering, University of Sydney(悉尼大学电气与计算机工程学院) School of Computer Science, University of Adelaide(阿德莱德大学计算机科学学院) School of Computing and Information Technology, University of Wollongong(沃林根大学计算与信息科技学院)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15220 2025-08-22 cs.LG cs.AI cs.LO 62%

Locally Pareto-Optimal Interpretations for Black-Box Machine Learning Models

Aniruddha Joshi, Supratik Chakraborty, S Akshay, Shetal Shah, Hazem Torfah, Sanjit Seshia

机构 * University of California at Berkeley(加州大学伯克利分校) Indian Institute of Technology Bombay(印度班加罗尔理工学院) Chalmers University of Technology and University of Gothenburg(查尔姆斯理工大学和哥德堡大学)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG

Comments This work has been accepted at ATVA'25

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.14921 2025-08-22 cs.CY cs.AI 62%

Designing an Interdisciplinary Artificial Intelligence Curriculum for Engineering: Evaluation and Insights from Experts

Johannes Schleiss, Anke Manukjan, Michelle Ines Bieber, Sebastian Lang, Sebastian Stober

机构 * Otto von Guericke University Magdeburg(奥托·冯·格里克大学马格德堡)

专题命中 安全评测 :alignment(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.04742 2025-08-22 cs.CY cs.AI 62%

A Case for Specialisation in Non-Human Entities

El-Mahdi El-Mhamdi, Lê-Nguyên Hoang, Mariame Tighanimine

专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.CY

Comments Accepted to AAAI/ACM AIES 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.14716 2025-08-22 cs.LG 57%

Pairwise or Pointwise? Evaluating Feedback Protocols for Bias in LLM-Based Evaluation

Tuhina Tripathi, Manya Wadhwa, Greg Durrett, Scott Niekum

专题命中 安全评测 :alignment(abstract);分类 cs.LG

Comments Published at COLM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15251 2025-08-22 eess.IV cs.AI cs.CV 57%

Explainable Knowledge Distillation for Efficient Medical Image Classification

Aqib Nazir Mir, Danish Raza Rizvi

机构 * Dept. of Computer Engineering(计算机工程系) Jamia Millia Islamia(Jamia Millia Islamia大学)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.14927 2025-08-22 cs.GT cs.AI 57%

AI Testing Should Account for Sophisticated Strategic Behaviour

Vojtech Kovarik, Eric Olav Chen, Sami Petersen, Alexis Ghersengorin, Vincent Conitzer

机构 * Department of Computer Science(计算机科学系) Czech Technical University Prague(捷克技术大学布拉格) Global Priorities Institute(全球优先研究所) University of Oxford(牛津大学) Foundations of Cooperative AI Lab(合作人工智能基础实验室) Carnegie Mellon University(卡内基梅隆大学)

专题命中 安全评测 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12259 2025-08-22 cs.CR cs.AI cs.ET 57%

Fortifying the Agentic Web: A Unified Zero-Trust Architecture Against Logic-layer Threats

Ken Huang, Yasir Mehmood, Hammad Atta, Jerry Huang, Muhammad Zeeshan Baig, Sree Bhargavi Balija

机构 * Qorvex Consulting Kleiner Perkins Wentworth Institute of Higher Education

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏