arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-09-24 至 2025-09-24 共收录 53 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 偏好对齐 8 篇

2501.04561 2025-09-24 cs.CL cs.CV 79%

OpenOmni: Advancing Open-Source Omnimodal Large Language Models with Progressive Multimodal Alignment and Real-Time Self-Aware Emotional Speech Synthesis

Run Luo, Ting-En Lin, Haonan Zhang, Yuchuan Wu, Xiong Liu, Min Yang, Yongbin Li, Longze Chen, Jiaming Li, Lei Zhang, Xiaobo Xia, Hamid Alinejad-Rokny, Fei Huang

机构 * Shenzhen Key Laboratory for High Performance Data Mining(深圳高性能数据挖掘重点实验室) Shenzhen Institute of Advanced Technology(深圳先进技术研究院) Chinese Academy of Sciences(中国科学院) University of Chinese Academy of Sciences(中国科学院大学) Tongyi Laboratory(通义实验室) University of New South Wales(新南威尔士大学) National University of Singapore(新加坡国立大学) University of Science and Technology of China(中国科学技术大学) MoE Key Laboratory of Brain-inspired Intelligent Perception and Cognition(脑启发智能感知与认知重点实验室)

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15044 2025-09-24 cs.CL 77%

Reward-Shifted Speculative Sampling Is An Efficient Test-Time Weak-to-Strong Aligner

Bolian Li, Yanran Wu, Xinyu Luo, Ruqi Zhang

机构 * Department of Computer Science, Purdue University(计算机科学系,普渡大学)

专题命中 偏好对齐 :alignment(abstract);RLHF(abstract);safety(abstract);分类 cs.CL

Comments EMNLP 2025 Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.24846 2025-09-24 cs.AI cs.CL 73%

MiCRo: Mixture Modeling and Context-aware Routing for Personalized Preference Learning

Jingyan Shen, Jiarui Yao, Rui Yang, Yifan Sun, Feng Luo, Rui Pan, Tong Zhang, Han Zhao

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) New York University(纽约大学) Rice University(稻谷大学)

专题命中 偏好对齐 :alignment(abstract);RLHF(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18632 2025-09-24 cs.CL 70%

A Good Plan is Hard to Find: Aligning Models with Preferences is Misaligned with What Helps Users

Nishant Balepur, Matthew Shu, Yoo Yeon Sung, Seraphina Goldfarb-Tarrant, Shi Feng, Fumeng Yang, Rachel Rudinger, Jordan Lee Boyd-Graber

专题命中 偏好对齐 :alignment(abstract);RLHF(abstract);分类 cs.CL

Comments EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19265 2025-09-24 cs.AI cs.CL 62%

Cross-Cultural Transfer of Commonsense Reasoning in LLMs: Evidence from the Arab World

Saeed Almheiri, Rania Hossam, Mena Attia, Chenxi Wang, Preslav Nakov, Timothy Baldwin, Fajri Koto

机构 * Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)

专题命中 偏好对齐 :alignment(abstract);分类 cs.CL、cs.AI

Comments EMNLP 2025 - Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18316 2025-09-24 cs.CL cs.AI 62%

Brittleness and Promise: Knowledge Graph Based Reward Modeling for Diagnostic Reasoning

Saksham Khatwani, He Cheng, Majid Afshar, Dmitriy Dligach, Yanjun Gao

机构 * University of Colorado Boulder(科罗拉多大学博尔德分校) University of Colorado Anschutz(科罗拉多大学安舒茨分校) University of Wisconsin - Madison(威斯康星大学麦迪逊分校) Loyola University Chicago(芝加哥洛约拉大学)

专题命中 偏好对齐 :trustworthy(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18661 2025-09-24 cs.IR cs.CL cs.HC 57%

Agentic AutoSurvey: Let LLMs Survey LLMs

Yixin Liu, Yonghui Wu, Denghui Zhang, Lichao Sun

机构 * Lehigh University(莱文大学) University of Florida(佛罗里达大学) Stevens Institute of Technology(史蒂文斯理工学院)

专题命中 偏好对齐 :RLHF(abstract);分类 cs.CL

Comments 29 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.14487 2025-09-24 cs.CV 50%

Token Preference Optimization with Self-Calibrated Visual-Anchored Rewards for Hallucination Mitigation

Jihao Gu, Yingyao Wang, Meng Cao, Pi Bu, Jun Song, Yancheng He, Shilong Li, Bo Zheng

机构 * Taobao & Tmall Group of Alibaba(淘宝与天猫集团)

专题命中 偏好对齐 :DPO(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 安全训练 8 篇

2509.19212 2025-09-24 cs.CL cs.AI 84%

Steering Multimodal Large Language Models Decoding for Context-Aware Safety

Zheyuan Liu, Zhangchen Xu, Guangyao Dou, Xiangchi Yuan, Zhaoxuan Tan, Radha Poovendran, Meng Jiang

机构 * University of Notre Dame(notre dame 大学) University of Washington(华盛顿大学) Johns Hopkins University(约翰霍普金斯大学) Georgia Institute of Technology(佐治亚理工学院)

专题命中 安全训练 :safety(title,abstract);alignment(abstract);分类 cs.CL、cs.AI

Comments A lightweight and model-agnostic decoding framework that dynamically adjusts token generation based on multimodal context

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18382 2025-09-24 cs.AI 79%

Evaluating the Safety and Skill Reasoning of Large Reasoning Models Under Compute Constraints

Adarsha Balaji, Le Chen, Rajeev Thakur, Franck Cappello, Sandeep Madireddy

机构 * Argonne National Laboratory(阿贡国家实验室) Mathematics and Computer Science Division(数学与计算机科学 division) Data Science and Learning Division(数据科学与学习 division)

专题命中 安全训练 :safety(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02159 2025-09-24 cs.LG 70%

PIGDreamer: Privileged Information Guided World Models for Safe Partially Observable Reinforcement Learning

Dongchi Huang, Jiaqi Wang, Yang Li, Chunhe Xia, Tianle Zhang, Kaige Zhang

机构 * School of Computer Science, University of Beihang, Beijing, China(北京航空航天大学计算机学院) School of Computer Science, Chinese University of Hong Kong, Hongkong, China(香港中文大学计算机学院) JD Explore Academy, Beijing, China(京东探索研究院) North Automatic Control Institute, Taiyuan, China(太原北自动控制研究所)

专题命中 安全训练 :alignment(abstract);safety(abstract);分类 cs.LG

Comments ICML 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.13255 2025-09-24 cs.CL cs.AI cs.IR cs.LG cs.MM 67%

Automating Steering for Safe Multimodal Large Language Models

Lyucheng Wu, Mengru Wang, Ziwen Xu, Tri Cao, Nay Oo, Bryan Hooi, Shumin Deng

机构 * Zhejiang University(浙江大学) Zhejiang University - Ant Group Joint Lab of Knowledge Graph(浙江大学-蚂蚁集团知识图谱联合实验室) National University of Singapore, NUS-NCS Joint Lab(新加坡国立大学NUS-NCS联合实验室)

专题命中 安全训练 :safety(abstract);分类 cs.CL、cs.AI、cs.LG

Comments EMNLP 2025 Main Conference. 23 pages (8+ for main); 25 figures; 1 table

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.03293 2025-09-24 cs.AI cs.CL cs.LG cs.SY eess.SY 67%

LogicGuard: Improving Embodied LLM agents through Temporal Logic based Critics

Anand Gokhale, Vaibhav Srivastava, Francesco Bullo

机构 * Department of Mechanical Engineering, University of California at Santa Barbara(加州大学圣芭芭拉分校机械工程系) Department of Electrical and Computer Engineering, Michigan State University(密歇根州立大学电气与计算机工程系)

专题命中 安全训练 :safety(abstract);分类 cs.CL、cs.AI、cs.LG

Comments Modified version of prior LTLCrit work with new robotics dataset

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11044 2025-09-24 cs.LG cs.AI q-bio.BM 62%

FragmentGPT: A Unified GPT Model for Fragment Growing, Linking, and Merging in Molecular Design

Xuefeng Liu, Songhao Jiang, Qinan Huang, Tinson Xu, Ian Foster, Mengdi Wang, Hening Lin, Rick Stevens

机构 * Department of Computer Science, University of Chicago(芝加哥大学计算机科学系) Pritzker School of Molecular Engineering, University of Chicago(芝加哥大学分子工程学院) Department of Chemistry, University of Chicago(芝加哥大学化学系) Argonne National Laboratory(阿贡国家实验室) AI Lab, Princeton University(普林斯顿大学人工智能实验室)

专题命中 安全训练 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.15389 2025-09-24 cs.CL cs.CR cs.CV 57%

Are Vision-Language Models Safe in the Wild? A Meme-Based Benchmark Study

DongGeon Lee, Joonwon Jang, Jihae Jeong, Hwanjo Yu

机构 * Pohang University of Science and Technology (POSTECH)(釜山科学技术大学)

专题命中 安全训练 :safety(abstract);分类 cs.CL

Comments Accepted to EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18681 2025-09-24 cs.AI 57%

Implementation of airborne ML models with semantics preservation

Nicolas Valot, Louis Fabre, Benjamin Lesage, Ammar Mechouche, Claire Pagetti

专题命中 安全训练 :safety(abstract);分类 cs.AI

Journal ref 44th Digital Avionics Systems Conference (DASC), Sep 2025, Montreal, Canada

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 越狱攻击 4 篇

2509.18058 2025-09-24 cs.LG cs.AI cs.CR 91%

Strategic Dishonesty Can Undermine AI Safety Evaluations of Frontier LLMs

Alexander Panfilov, Evgenii Kortukov, Kristina Nikolić, Matthias Bethge, Sebastian Lapuschkin, Wojciech Samek, Ameya Prabhu, Maksym Andriushchenko, Jonas Geiping

机构 * ELLIS Institute Tübingen & MPI for Intelligent Systems(图宾根ELLIS研究所与智能系统马克斯·普朗克研究所) Tübingen AI Center(图宾根人工智能中心) Fraunhofer HHI(弗劳恩霍夫高精尖研究所) ETH Zurich & ETH AI Center(苏黎世联邦理工学院与苏黎世人工智能中心) University of Tübingen(图宾根大学) TU Dublin(都柏林技术大学) TU Berlin & BIFOLD(柏林技术大学与BIFOLD)

专题命中 越狱攻击 :safety(title,abstract);AI safety(title);alignment(abstract);jailbreak(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.13329 2025-09-24 cs.CL cs.AI cs.LG 67%

Language Models Can Predict Their Own Behavior

Dhananjay Ashok, Jonathan May

机构 * Information Sciences Institute(信息科学研究所) University of Southern California(南加州大学)

专题命中 越狱攻击 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments Presented at the Thirty-Ninth Annual Conference on Neural Information Processing Systems (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19100 2025-09-24 cs.LG cs.AI 62%

Algorithms for Adversarially Robust Deep Learning

Alexander Robey

专题命中 越狱攻击 :safety(abstract);分类 cs.AI、cs.LG

Comments PhD thesis

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17371 2025-09-24 cs.CR cs.LG 57%

SilentStriker:Toward Stealthy Bit-Flip Attacks on Large Language Models

Haotian Xu, Qingsong Peng, Jie Shi, Huadi Zheng, Yu Li, Cheng Zhuo

机构 * Zhejiang University(浙江大学) Huawei(华为)

专题命中 越狱攻击 :prompt injection(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 幻觉与事实性 3 篇

2509.18792 2025-09-24 cs.CL 57%

Beyond the Leaderboard: Understanding Performance Disparities in Large Language Models via Model Diffing

Sabri Boughorbel, Fahim Dalvi, Nadir Durrani, Majd Hawasly

机构 * Qatar Computing Research Institute, HBKU(卡塔尔计算研究所,哈瓦那大学)

专题命中 幻觉与事实性 :safety(abstract);分类 cs.CL

Comments 12 pages, accepted to the 2025 Conference on Empirical Methods in Natural Language Processing (EMNLP 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18132 2025-09-24 cs.AI 57%

Position Paper: Integrating Explainability and Uncertainty Estimation in Medical AI

Xiuyi Fan

机构 * Lee Kong Chian School of Medicine, College of Computing Data Science, Nanyang Technological University, Singapore

专题命中 幻觉与事实性 :trustworthy(abstract);分类 cs.AI

Comments Accepted at the International Joint Conference on Neural Networks, IJCNN 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18128 2025-09-24 cs.LG 57%

Accounting for Uncertainty in Machine Learning Surrogates: A Gauss-Hermite Quadrature Approach to Reliability Analysis

Amirreza Tootchi, Xiaoping Du

专题命中 幻觉与事实性 :trustworthy(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 隐私与版权 3 篇

2509.18104 2025-09-24 cs.LG cs.AI 62%

Data Valuation and Selection in a Federated Model Marketplace

Wenqian Li, Youjia Yang, Ruoxi Jia, Yan Pang

机构 * IORA National University of Singapore(国际研究机构国立新加坡大学) NUS Research Institution(国立新加坡大学研究机构) Department of ECE(电子工程系) Virginia Tech(弗吉尼亚理工大学) Business School(商学院)

专题命中 隐私与版权 :trustworthy(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18400 2025-09-24 cs.AI 57%

ATLAS: Benchmarking and Adapting LLMs for Global Trade via Harmonized Tariff Code Classification

Pritish Yuvraj, Siva Devarakonda

专题命中 隐私与版权 :alignment(abstract);分类 cs.AI

Journal ref Paper in Review For ICLR 2026 (Workshop)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19041 2025-09-24 cs.HC 50%

Position: Human-Robot Interaction in Embodied Intelligence Demands a Shift From Static Privacy Controls to Dynamic Learning

Shuning Zhang, Hong Jia, Simin Li, Ting Dang, Yongquan `Owen' Hu, Xin Yi, Hewu Li

专题命中 隐私与版权 :trustworthy(abstract)

Comments To be published in NeurIPS 2025 Workshop on Bridging Language, Agent, and World Models for Reasoning and Planning

详情

展开后加载摘要…

URL PDF HTML 收藏

6. 安全评测 10 篇

2501.01346 2025-09-24 cs.CV cs.CL 79%

Large Vision-Language Model Alignment and Misalignment: A Survey Through the Lens of Explainability

Dong Shu, Haiyan Zhao, Jingyu Hu, Weiru Liu, Ali Payani, Lu Cheng, Mengnan Du

机构 * Northwestern University(西北大学) New Jersey Institute of Technology(新泽西理工学院) University of Bristol(布里斯托大学) Cisco Research(思科研究) University of Illinois Chicago(伊利诺伊大学芝加哥分校)

专题命中 安全评测 :alignment(title,abstract);分类 cs.CL

Comments EMNLP 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15260 2025-09-24 cs.CL 79%

Toxicity Red-Teaming: Benchmarking LLM Safety in Singapore's Low-Resource Languages

Yujia Hu, Ming Shan Hee, Preslav Nakov, Roy Ka-Wei Lee

机构 * Singapore University of Technology and Design(新加坡科技设计大学) Mohamed bin Zayed University of Artificial Intelligence(马尔代夫 bin Zayed 人工智能大学)

专题命中 安全评测 :safety(title,abstract);分类 cs.CL

Comments 9 pages, EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19120 2025-09-24 cs.LG cs.AI cs.DC 76%

FedFiTS: Fitness-Selected, Slotted Client Scheduling for Trustworthy Federated Learning in Healthcare AI

Ferdinand Kahenga, Antoine Bagula, Sajal K. Das, Patrick Sello

机构 * Department of Computer Science University of the Western Cape(计算机科学系,西开普敦大学) Department of Computer Science Missouri University of Science and Technology(计算机科学系,密苏里科学与技术大学)

专题命中 安全评测 :trustworthy(title);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18557 2025-09-24 cs.AI 70%

LLMZ+: Contextual Prompt Whitelist Principles for Agentic LLMs

Tom Pawelek, Raj Patel, Charlotte Crowell, Noorbakhsh Amiri, Sudip Mittal, Shahram Rahimi, Andy Perkins

机构 * Department of Computer Science Mississippi State University(计算机科学系密苏里州立大学) Department of Computer Science The University of Alabama(计算机科学系阿拉巴马大学) Mississippi State University, Mississippi State, MS, USA(密苏里州立大学) The University of Alabama, Tuscaloosa, AL, USA(阿拉巴马大学)

专题命中 安全评测 :jailbreak(abstract);prompt injection(abstract);分类 cs.AI

Comments 7 pages, 5 figures, to be published and presented at ICMLA 2025

详情

展开后加载摘要…

URL PDF HTML 收藏