arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-10-13 至 2025-10-13 共收录 43 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 偏好对齐 4 篇

2506.03517 2025-10-13 cs.CV 67%

DenseDPO: Fine-Grained Temporal Preference Optimization for Video Diffusion Models

Ziyi Wu, Anil Kag, Ivan Skorokhodov, Willi Menapace, Ashkan Mirzaei, Igor Gilitschenski, Sergey Tulyakov, Aliaksandr Siarohin

机构 * Snap Research University of Toronto(多伦多大学) Vector Institute(向量研究所)

专题命中 偏好对齐 :alignment(abstract);DPO(abstract)

Comments NeurIPS 2025 Spotlight. Project page: https://snap-research.github.io/DenseDPO/

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.09103 2025-10-13 cs.LG 57%

AdaPM: a Partial Momentum Algorithm for LLM Training

Yimu Zhang, Yuanshi Liu, Cong Fang

机构 * State Key Lab of General AI, School of Intelligence Science and Technology, Peking University(通用人工智能国家重点实验室,智能科学与技术学院,北京大学) Institute for Artificial Intelligence, Peking University(人工智能研究院,北京大学)

专题命中 偏好对齐 :RLHF(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.13726 2025-10-13 cs.CL 57%

RPO: Retrieval Preference Optimization for Robust Retrieval-Augmented Generation

Shi-Qi Yan, Quan Liu, Zhen-Hua Ling

机构 * National Engineering Research Center of Speech and Language Information Processing, University of Science and Technology of China(语音与语言信息处理国家工程研究中心,中国科学技术大学) State Key Laboratory of Cognitive Intelligence, iFLYTEK Research(认知智能国家重点实验室,iFLYTEK研究院)

专题命中 偏好对齐 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08912 2025-10-13 cs.HC 50%

Beyond Words: Infusing Conversational Agents with Human-like Typing Behaviors

Jijie Zhou, Yuhan Hu

专题命中 偏好对齐 :trustworthy(abstract)

Comments Author's version of a paper published at CUI '24 (ACM Conversational User Interfaces 2024)

Journal ref CUI '24: Proceedings of the ACM Conversational User Interfaces 2024, July 8-10, 2024, Luxembourg, Luxembourg. ACM, New York, NY, USA, 11 pages

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 安全训练 10 篇

2510.09004 2025-10-13 cs.CL 89%

Decoupling Safety into Orthogonal Subspace: Cost-Efficient and Performance-Preserving Alignment for Large Language Models

Yutao Mou, Xiaoling Zhou, Yuxiao Luo, Shikun Zhang, Wei Ye

机构 * National Engineering Research Center for Software Engineering, Peking University, China(软件工程国家工程研究中心,北京大学,中国)

专题命中 安全训练 :alignment(title,abstract);safety(title,abstract);trustworthy(abstract);分类 cs.CL

Comments Work in Progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06944 2025-10-13 cs.LG cs.AI cs.CL cs.CV 67%

AMFT: Aligning LLM Reasoners by Meta-Learning the Optimal Imitation-Exploration Balance

Lixuan He, Jie Feng, Yong Li

专题命中 安全训练 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments The paper is currently under investigation regarding concerns of potential academic misconduct. While the investigation is ongoing, the authors have voluntarily requested to withdraw the manuscript

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08776 2025-10-13 cs.CL cs.AI 62%

Measuring Moral LLM Responses in Multilingual Capacities

Kimaya Basu, Savi Kolari, Allison Yu

专题命中 安全训练 :safety(abstract);分类 cs.CL、cs.AI

Comments 10 pages, 5 figures; referenced articles: arXiv:2303.08774, arXiv:2303.12528, arXiv:2308.14132, arXiv:2505.12201, arXiv:2406.04428, arXiv:2407.02273, arXiv:2404.01268, arXiv:2502.09747, arXiv:2507.13474, arXiv:2505.21479, arXiv:2306.05685

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08648 2025-10-13 cs.LG cs.AI 62%

Inverse-Free Wilson Loops for Transformers: A Practical Diagnostic for Invariance and Order Sensitivity

Edward Y. Chang, Ethan Y. Chang

专题命中 安全训练 :safety(abstract);分类 cs.AI、cs.LG

Comments 24 pages, 10 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04503 2025-10-13 cs.CR cs.AI cs.CL 62%

P2P: A Poison-to-Poison Remedy for Reliable Backdoor Defense in LLMs

Shuai Zhao, Xinyi Wu, Shiqian Zhao, Xiaobao Wu, Zhongliang Guo, Yanhao Jia, Anh Tuan Luu

机构 * Nanyang Technological University, Singapore(南洋理工大学) Shanghai Jiao Tong University, Shanghai, China(上海交通大学)

专题命中 安全训练 :trustworthy(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.11411 2025-10-13 cs.LG 57%

Detecting and Filtering Unsafe Training Data via Data Attribution with Denoised Representation

Yijun Pan, Taiwei Shi, Jieyu Zhao, Jiaqi W. Ma

机构 * University of Michigan(密歇根大学) University of Southern California(南加州大学) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

专题命中 安全训练 :trustworthy(abstract);分类 cs.LG

Comments 14 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.08428 2025-10-13 cs.RO cs.AI cs.SY eess.SY 57%

SwarmGPT: Combining Large Language Models with Safe Motion Planning for Drone Swarm Choreography

Martin Schuck, Dinushka Orrin Dahanaggamaarachchi, Ben Sprenger, Vedant Vyas, Siqi Zhou, Angela P. Schoellig

机构 * Learning Systems and Robotics Lab(学习系统与机器人实验室) Munich Institute of Robotics and Machine Intelligence(慕尼黑机器人与机器智能研究所) Technical University of Munich(慕尼黑技术大学)

专题命中 安全训练 :safety(abstract);分类 cs.AI

Comments Accepted at RA-L 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.09200 2025-10-13 cs.CV cs.AI cs.HC 57%

Towards Safer and Understandable Driver Intention Prediction

Mukilan Karuppasamy, Shankar Gangisetty, Shyam Nandan Rai, Carlo Masone, C V Jawahar

机构 * IIIT Hyderabad(海得拉巴印度理工学院) Politecnico di Torino(托里诺理工学院)

专题命中 安全训练 :safety(abstract);分类 cs.AI

Comments 10 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.13336 2025-10-13 cs.RO cs.LG 57%

Maximizing UAV Cellular Connectivity with Reinforcement Learning for BVLoS Path Planning

Mehran Behjati, Rosdiadee Nordin, Nor Fadzilah Abdullah

专题命中 安全训练 :safety(abstract);分类 cs.LG

Comments Submitted to an IEEE Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08917 2025-10-13 cs.HC 50%

"I know it's not right, but that's what it said to do": Investigating Trust in AI Chatbots for Cybersecurity Policy

Brandon Lit, Edward Crowder, Daniel Vogel, Hassan Khan

专题命中 安全训练 :prompt injection(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 越狱攻击 4 篇

2510.09023 2025-10-13 cs.LG cs.CR 79%

The Attacker Moves Second: Stronger Adaptive Attacks Bypass Defenses Against Llm Jailbreaks and Prompt Injections

Milad Nasr, Nicholas Carlini, Chawin Sitawarin, Sander V. Schulhoff, Jamie Hayes, Michael Ilie, Juliette Pluto, Shuang Song, Harsh Chaudhari, Ilia Shumailov, Abhradeep Thakurta, Kai Yuanqing Xiao, Andreas Terzis, Florian Tramèr

机构 * OpenAI Anthropic Google DeepMind(谷歌DeepMind) HackAPrompt Northeastern University(东北大学) ETH Zürich(苏黎世联邦理工学院) AI Sequrity Company(AI安全公司) MATS Main contributors(MATS主要贡献者)

专题命中 越狱攻击 :prompt injection(title,abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.06594 2025-10-13 cs.CL 79%

Do Internal Layers of LLMs Reveal Patterns for Jailbreak Detection?

Sri Durga Sai Sowmya Kadali, Evangelos E. Papalexakis

机构 * Dept. of Computer Science and Engineering University of California, Riverside(计算机科学与工程系加州大学河滨分校)

专题命中 越狱攻击 :jailbreak(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01644 2025-10-13 cs.CL cs.AI cs.CY 75%

Machine Learning for Detection and Analysis of Novel LLM Jailbreaks

John Hawkins, Aditya Pramar, Rodney Beard, Rohitash Chandra

机构 * Centre for Artificial Intelligence and Innovation(人工智能与创新中心) Pingla Institute(平拉研究所) Transitional Artificial Intelligence Research Group(过渡人工智能研究组) UNSW(新南威尔士大学)

专题命中 越狱攻击 :safety(abstract);jailbreak(abstract);分类 cs.CL、cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.09471 2025-10-13 cs.CL 70%

Getting Your Indices in a Row: Full-Text Search for LLM Training Data for Real World

Ines Altemir Marinas, Anastasiia Kucherenko, Alexander Sternfeld, Andrei Kucharavy

机构 * IC, EPFL(EPFL)

专题命中 越狱攻击 :safety(abstract);jailbreak(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 提示注入 1 篇

2510.08829 2025-10-13 cs.CR cs.AI cs.LG 73%

CommandSans: Securing AI Agents with Surgical Precision Prompt Sanitization

Debeshee Das, Luca Beurer-Kellner, Marc Fischer, Maximilian Baader

机构 * ETH Zurich(苏黎世联邦理工学院) Snyk(Snyk公司)

专题命中 提示注入 :safety(abstract);prompt injection(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 幻觉与事实性 3 篇

2505.21641 2025-10-13 cs.LG cs.CR stat.ME 57%

PrivATE: Differentially Private Confidence Intervals for Average Treatment Effects

Maresa Schröder, Justin Hartenstein, Stefan Feuerriegel

机构 * LMU Munich(慕尼黑莱茵河大学) Munich Center for Machine Learning(慕尼黑机器学习中心) Stanford University(斯坦福大学)

专题命中 幻觉与事实性 :safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.12444 2025-10-13 cs.CL 57%

Augmenting Compliance-Guaranteed Customer Service Chatbots: Context-Aware Knowledge Expansion with Large Language Models

Mengze Hong, Chen Jason Zhang, Di Jiang, Yuanqin He

机构 * Hong Kong Polytechnic University(香港理工大学) AI Group, WeBank Co., Ltd(WeBank金融科技部)

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.CL

Comments Accepted by EMNLP 2025 Industry Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08658 2025-10-13 econ.GN q-fin.EC 50%

When Truth Does Not Take on Its Shoes: How Misinformation Spreads in Chatrooms

Shuige Liu

专题命中 幻觉与事实性 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

6. 隐私与版权 1 篇

2510.09434 2025-10-13 cs.CL 57%

Domain-Adapted Pre-trained Language Models for Implicit Information Extraction in Crash Narratives

Xixi Wang, Jordanka Kovaceva, Miguel Costa, Shuai Wang, Francisco Camara Pereira, Robert Thomson

机构 * Department of Technology, Management and Economics, Technical University of Denmark(技术、管理与经济系,丹麦技术大学) Department of Mechanics and Maritime Sciences, Chalmers University of Technology(机械与航海科学系,查尔姆斯理工大学) Department of Computer Science and Engineering, Chalmers University of Technology(计算机科学与工程系,查尔姆斯理工大学)

专题命中 隐私与版权 :safety(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏

7. 安全评测 6 篇

2407.08983 2025-10-13 cs.SE cs.AI cs.LG 81%

Towards More Trustworthy and Interpretable LLMs for Code through Syntax-Grounded Explanations

David N. Palacio, Daniel Rodriguez-Cardenas, Alejandro Velasco, Dipin Khati, Kevin Moran, Denys Poshyvanyk

机构 * University of Central Florida(佛罗里达中央大学)

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI、cs.LG

Comments Under Review to appear in ACM Transactions on Software Engineering and Methodology (TOSEM)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.09090 2025-10-13 cs.CY cs.AI 76%

AI and Human Oversight: A Risk-Based Framework for Alignment

Laxmiraju Kandikatla, Branislav Radeljic

专题命中 安全评测 :alignment(title);分类 cs.AI、cs.CY

Comments 19 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08608 2025-10-13 cs.CL cs.AI 76%

MMA-ASIA: A Multilingual and Multimodal Alignment Framework for Culturally-Grounded Evaluation

Weihua Zheng, Zhengyuan Liu, Tanmoy Chakraborty, Weiwen Xu, Xiaoxue Gao, Bryan Chen Zhengyu Tan, Bowei Zou, Chang Liu, Yujia Hu, Xing Xie, Xiaoyuan Yi, Jing Yao, Chaojun Wang, Long Li, Rui Liu, Huiyao Liu, Koji Inoue, Ryuichi Sumida, Tatsuya Kawahara, Fan Xu, Lingyu Ye, Wei Tian, Dongjun Kim, Jimin Jung, Jaehyung Seo, Nadya Yuki Wangsajaya, Pham Minh Duc, Ojasva Saxena, Palash Nandi, Xiyan Tao, Wiwik Karlina, Tuan Luong, Keertana Arun Vasan, Roy Ka-Wei Lee, Nancy F. Chen

机构 * Singapore University of Technology and Design(新加坡科技设计大学) Agency for Science, Technology and Research, Singapore(新加坡科技研究局) Indian Institute of Technology Delhi(印度理工学院德里分校) Alibaba DAMO Academy(阿里巴巴达摩院) Microsoft Research Asia(微软亚洲研究院) Shanghai University of Finance and Economics(上海财经大学) Inner Mongolia University(内蒙古大学) Kyoto University(京都大学) Jiangxi Normal University(江西师范大学) Korea University(韩国大学) Nanyang Technological University(南洋理工大学)

专题命中 安全评测 :alignment(title);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08930 2025-10-13 cs.HC cs.AI 57%

Co-Authoring the Self: A Human-AI Interface for Interest Reflection in Recommenders

Ruixuan Sun, Junyuan Wang, Sanjali Roy, Joseph A. Konstan

机构 * Grouplens Research, University of Minnesota(Grouplens研究院、明尼苏达大学) Department of Computer Science and Engineering, University of Minnesota(计算机科学与工程系、明尼苏达大学)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08783 2025-10-13 cs.HC cs.AI 57%

MLLM as a UI Judge: Benchmarking Multimodal LLMs for Predicting Human Perception of User Interfaces

Reuben A. Luera, Ryan Rossi, Franck Dernoncourt, Samyadeep Basu, Sungchul Kim, Subhojyoti Mukherjee, Puneet Mathur, Ruiyi Zhang, Jihyung Kil, Nedim Lipka, Seunghyun Yoon, Jiuxiang Gu, Zichao Wang, Cindy Xiong Bearfield, Branislav Kveton

机构 * University of California, Berkeley(加州大学伯克利分校) Adobe Research(Adobe研究) Georgia Institute of Technology(佐治亚理工学院)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.09320 2025-10-13 cs.CV 50%

Hybrid-grained Feature Aggregation with Coarse-to-fine Language Guidance for Self-supervised Monocular Depth Estimation

Wenyao Zhang, Hongsi Liu, Bohan Li, Jiawei He, Zekun Qi, Yunnan Wang, Shengyang Zhao, Xinqiang Yu, Wenjun Zeng, Xin Jin

机构 * MoE Key Lab of Artificial Intelligence, AI Institute, Shanghai Jiao Tong University(人工智能 MOE 实验室,上海交通大学人工智能学院) Ningbo Institute of Digital Twin, Eastern Institute of Technology, Ningbo, China(宁波数字孪生研究所,东技术研究所,宁波,中国) Ningbo Key Laboratory of Spatial Intelligence and Digital Derivative, Ningbo, China(宁波空间智能与数字衍生关键实验室,宁波,中国) University of Science and Technology of China(中国科学技术大学) CASIA Tsinghua University(清华大学)

专题命中 安全评测 :alignment(abstract)

Comments ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏

8. AI治理与伦理 1 篇

2510.09051 2025-10-13 cs.CL cs.AI cs.LG 67%

Alif: Advancing Urdu Large Language Models via Multilingual Synthetic Data Distillation

Muhammad Ali Shafique, Kanwal Mehreen, Muhammad Arham, Maaz Amjad, Sabur Butt, Hamza Farooq

机构 * University of British Columbia(不列颠哥伦比亚大学) Texas Tech University(德克萨斯技术大学) Institute for the Future of Education, Tecnológico de Monterrey(教育未来研究所,墨西哥蒙特雷技术学院)

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CL、cs.AI、cs.LG

Comments Accepted to the EMNLP 2025 Workshop on Multilingual Representation Learning (MRL)

详情

展开后加载摘要…

URL PDF HTML 收藏