arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-09-26 至 2025-09-26 共收录 40 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 偏好对齐 8 篇

2508.07137 2025-09-26 cs.LG cs.AI 86%

A Principled Loss Function for Direct Language Model Alignment

Yuandong Tan

专题命中 偏好对齐 :alignment(title,abstract);RLHF(abstract);DPO(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.20938 2025-09-26 cs.RO cs.CV 82%

Autoregressive End-to-End Planning with Time-Invariant Spatial Alignment and Multi-Objective Policy Refinement

Jianbo Zhao, Taiyu Ban, Xiangjie Li, Xingtai Gui, Hangning Zhou, Lei Liu, Hongwei Zhao, Bin Li

机构 * University of Science and Technology of China(中国科学技术大学) Mach Drive(马车驱动)

专题命中 偏好对齐 :alignment(title,abstract);DPO(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.14096 2025-09-26 cs.CV 78%

VideoPASTA: 7K Preference Pairs That Matter for Video-LLM Alignment

Yogesh Kulkarni, Pooyan Fazli

机构 * Arizona State University(亚利桑那州立大学)

专题命中 偏好对齐 :alignment(title,abstract)

Comments EMNLP 2025 (Main)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19249 2025-09-26 cs.CL cs.AI cs.LG 67%

Reinforcement Learning on Pre-Training Data

Siheng Li, Kejiao Li, Zenan Xu, Guanhua Huang, Evander Yang, Kun Li, Haoyuan Wu, Jiajia Wu, Zihao Zheng, Chenchen Zhang, Kun Shi, Kyrierl Deng, Qi Yi, Ruibin Xiong, Tingqiang Xu, Yuhao Jiang, Jianfeng Yan, Yuyuan Zeng, Guanghui Xu, Jinbao Xue, Zhijiang Xu, Zheng Fang, Shuai Li, Qibin Liu, Xiaoxue Li, Zhuoyu Li, Yangyu Tao, Fei Gao, Cheng Jiang, Bo Chao Wang, Kai Liu, Jianchen Zhu, Wai Lam, Wayyt Wang, Bo Zhou, Di Wang

机构 * LLM Department, Tencent(腾讯大模型部门) HunYuan Infra Team(文心一言基础设施团队) The Chinese University of Hong Kong(香港中文大学)

专题命中 偏好对齐 :RLHF(abstract);分类 cs.CL、cs.AI、cs.LG

Comments Work in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04897 2025-09-26 cs.CL cs.AI cs.LG 67%

PLaMo 2 Technical Report

Preferred Networks, :, Kaizaburo Chubachi, Yasuhiro Fujita, Shinichi Hemmi, Yuta Hirokawa, Kentaro Imajo, Toshiki Kataoka, Goro Kobayashi, Kenichi Maehashi, Calvin Metzger, Hiroaki Mikami, Shogo Murai, Daisuke Nishino, Kento Nozawa, Toru Ogawa, Shintarou Okada, Daisuke Okanohara, Shunta Saito, Shotaro Sano, Shuji Suzuki, Kuniyuki Takahashi, Daisuke Tanaka, Avinash Ummadisingu, Hanqin Wang, Sixue Wang, Tianqi Xu

机构 * Preferred Networks Inc(Preferred Networks公司)

专题命中 偏好对齐 :DPO(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.12148 2025-09-26 cs.CV 67%

HermesFlow: Seamlessly Closing the Gap in Multimodal Understanding and Generation

Ling Yang, Xinchen Zhang, Ye Tian, Chenming Shang, Minghao Xu, Wentao Zhang, Bin Cui

机构 * Peking University(北京大学) Tsinghua University(清华大学) Mila - Québec AI Institute(魁北克人工智能研究所)

专题命中 偏好对齐 :alignment(abstract);DPO(abstract)

Comments NeurIPS 2025. Code: https://github.com/Gen-Verse/HermesFlow

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.16663 2025-09-26 cs.CL cs.AI 62%

Turning Internal Gap into Self-Improvement: Promoting the Generation-Understanding Unification in MLLMs

Yujin Han, Hao Chen, Andi Han, Zhiheng Wang, Xinyu Liu, Yingya Zhang, Shiwei Zhang, Difan Zou

专题命中 偏好对齐 :DPO(abstract);分类 cs.CL、cs.AI

Comments 31 pages, 16 figures, 12 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.20961 2025-09-26 cs.CV cs.AI 57%

Unlocking Financial Insights: An advanced Multimodal Summarization with Multimodal Output Framework for Financial Advisory Videos

Sarmistha Das, R E Zera Marveen Lyngkhoi, Sriparna Saha, Alka Maurya

机构 * Indian Institute of Technology Patna(印度帕纳杰大学) CRISIL LTD(CRISIL公司)

专题命中 偏好对齐 :DPO(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 安全训练 6 篇

2509.20513 2025-09-26 cs.AI cs.DC 79%

Reconstruction-Based Adaptive Scheduling Using AI Inferences in Safety-Critical Systems

Samer Alshaer, Ala Khalifeh, Roman Obermaisser

专题命中 安全训练 :safety(title,abstract);分类 cs.AI

Comments 14 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.02781 2025-09-26 q-bio.QM cs.AI cs.LG 73%

Multimodal AI predicts clinical outcomes of drug combinations from preclinical data

Yepeng Huang, Xiaorui Su, Varun Ullanat, Intae Moon, Ivy Liang, Lindsay Clegg, Damilola Olabode, Ruthie Johnson, Nicholas Ho, Megan Gibbs, Megan Gibbs, Alexander Gusev, Bino John, Marinka Zitnik

机构 * Harvard Medical School(哈佛医学院) Harvard College(哈佛学院) AstraZeneca(阿斯利康) Carnegie Mellon University(卡内基梅隆大学) Dana-Farber Cancer Institute and Harvard Medical School(达纳-法伯癌症研究所和哈佛医学院) Harvard University(哈佛大学) Broad Institute of MIT and Harvard(MIT和哈佛大学 Broad研究所) Harvard Data Science Initiative(哈佛数据科学计划)

专题命中 安全训练 :alignment(abstract);safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.20792 2025-09-26 cs.CV cs.AI cs.LG 66%

DAC-LoRA: Dynamic Adversarial Curriculum for Efficient and Robust Few-Shot Adaptation

Ved Umrajkar

机构 * Indian Institute of Technology, Roorkee(印度理工学院拉胡尔分校)

专题命中 安全训练 :safety(abstract);分类 cs.AI、cs.LG;trustworthy(comments)

Comments Accepted at ICCV2025 Workshop on Safe and Trustworthy Multimodal AI Systems

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.07848 2025-09-26 cs.AI 57%

Safe Explicable Policy Search

Akkamahadevi Hanni, Jonathan Montaño, Yu Zhang

机构 * School of Computing and Augmented Intelligence, Arizona State University(计算与增强智能学院,亚利桑那州立大学) School of Mathematical and Statistical Sciences, Arizona State University(数学与统计科学学院,亚利桑那州立大学)

专题命中 安全训练 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.06123 2025-09-26 cs.LG cs.RO 57%

Security of Deep Reinforcement Learning for Autonomous Driving: A Survey

Ambra Demontis, Srishti Gupta, Maura Pintor, Luca Demetrio, Kathrin Grosse, Hsiao-Ying Lin, Chengfang Fang, Battista Biggio, Fabio Roli

专题命中 安全训练 :safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21102 2025-09-26 cs.CV 50%

Mammo-CLIP Dissect: A Framework for Analysing Mammography Concepts in Vision-Language Models

Suaiba Amina Salahuddin, Teresa Dorszewski, Marit Almenning Martiniussen, Tone Hovda, Antonio Portaluri, Solveig Thrun, Michael Kampffmeyer, Elisabeth Wetzer, Kristoffer Wickstrøm, Robert Jenssen

机构 * UiT The Arctic University of Norway(乌塔大学极地大学) Technical University of Denmark(技术大学) Østfold Hospital Trust(奥斯fold医院信托) Vestre Viken Hospital Trust(维斯特维肯医院信托) Radboud University Nijmegen Medical Centre(拉德堡德大学奈梅亨医疗中心) The Netherlands Cancer Institute(荷兰癌症研究所) Antoni van Leeuwenhoek Hospital(安东尼·弗莱明医院) University of Copenhagen(哥本哈根大学)

专题命中 安全训练 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 越狱攻击 1 篇

2411.00827 2025-09-26 cs.CV cs.AI 77%

IDEATOR: Jailbreaking and Benchmarking Large Vision-Language Models Using Themselves

Ruofan Wang, Juncheng Li, Yixu Wang, Bo Wang, Xiaosen Wang, Yan Teng, Yingchun Wang, Xingjun Ma, Yu-Gang Jiang

机构 * Fudan University(复旦大学) Huawei Technologies Ltd.(华为技术有限公司) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)

专题命中 越狱攻击 :alignment(abstract);safety(abstract);jailbreak(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 红队测试 1 篇

2509.21011 2025-09-26 cs.CR cs.AI cs.SE 79%

Automatic Red Teaming LLM-based Agents with Model Context Protocol Tools

Ping He, Changjiang Li, Binbin Zhao, Tianyu Du, Shouling Ji

机构 * College of Computer Science and Technology, Zhejiang University(浙江大学计算机科学与技术学院) Palo Alto Networks(帕洛阿尔托网络公司) School of Software Technology, Zhejiang University(浙江大学软件技术学院)

专题命中 红队测试 :red teaming(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 提示注入 1 篇

2509.10248 2025-09-26 cs.LG 79%

Prompt Injection Attacks on LLM Generated Reviews of Scientific Publications

Janis Keuper

机构 * Institute for Machine Learning and Analytics (IMLA)(机器学习与分析研究所)

专题命中 提示注入 :prompt injection(title,abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏

6. 安全评测 12 篇

2509.20393 2025-09-26 cs.CY cs.AI cs.LG 78%

The Secret Agenda: LLMs Strategically Lie and Our Current Safety Tools Are Blind

Caleb DeLeeuw, Gaurav Chawla, Aniket Sharma, Vanessa Dietze

机构 * Independent Researcher(独立研究者)

专题命中 安全评测 :safety(title);分类 cs.AI、cs.CY、cs.LG

Comments 9 pages plus citations and appendix, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.20680 2025-09-26 cs.LG cs.CL cs.CR 73%

Can Federated Learning Safeguard Private Data in LLM Training? Vulnerabilities, Attacks, and Defense Evaluation

Wenkai Guo, Xuefeng Liu, Haolin Wang, Jianwei Niu, Shaojie Tang, Jing Yuan

机构 * State Key Laboratory of Virtual Reality Technology and Systems, School of Computer Science and Engineering, Beihang University, Beijing, China(虚拟现实技术与系统国家重点实验室,计算机科学与工程学院,北京航空航天大学) Hangzhou Innovation Institute of Beihang University, Zhejiang Key Laboratory of Industrial Big Data and Robot Intelligent Systems, Hangzhou, China(北京航空航天大学杭州创新研究院,浙江省工业大数据与机器人智能系统重点实验室) Center for AI Business Innovation, Department of Management Science and Systems, University at Buffalo, Buffalo, New York, USA(人工智能商业创新中心,管理科学与系统系,布法罗大学) University of North Texas, Denton, Texas, USA(德克萨斯大学达文波特分校) Zhongguancun Laboratory, Beijing, China(中关村实验室)

专题命中 安全评测 :alignment(abstract);safety(abstract);分类 cs.CL、cs.LG

Comments 28 pages, 32 figures, accepted to the Findings of EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.20998 2025-09-26 cs.AI 70%

CORE: Full-Path Evaluation of LLM Agents Beyond Final State

Panagiotis Michelakis, Yiannis Hadjiyiannis, Dimitrios Stamoulis

机构 * Synkrasis Labs(Synkrasis实验室) Harbin Institute of Technology(哈尔滨工业大学)

专题命中 安全评测 :alignment(abstract);safety(abstract);分类 cs.AI

Comments Accepted: LAW 2025 Workshop NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21287 2025-09-26 cs.CL cs.AI 62%

DisCoCLIP: A Distributional Compositional Tensor Network Encoder for Vision-Language Understanding

Kin Ian Lo, Hala Hawashin, Mina Abbaszadeh, Tilen Limback-Stokin, Hadi Wazni, Mehrnoosh Sadrzadeh

机构 * University College London(伦敦大学学院)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.20520 2025-09-26 cs.AI cs.DC cs.LG 62%

Adaptive Approach to Enhance Machine Learning Scheduling Algorithms During Runtime Using Reinforcement Learning in Metascheduling Applications

Samer Alshaer, Ala Khalifeh, Roman Obermaisser

专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.LG

Comments 18 pages, 21 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.20378 2025-09-26 cs.CL cs.AI 62%

Beyond Global Emotion: Fine-Grained Emotional Speech Synthesis with Dynamic Word-Level Modulation

Sirui Wang, Andong Chen, Tiejun Zhao

机构 * Harbin Institute of Technology(哈尔滨工业大学)

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.11771 2025-09-26 cs.CL cs.AI 62%

The Validation Gap: A Mechanistic Analysis of How Language Models Compute Arithmetic but Fail to Validate It

Leonardo Bertolazzi, Philipp Mondorf, Barbara Plank, Raffaella Bernardi

机构 * DISI, University of Trento(特伦托大学DISI中心) MaiNLP, Center for Information and Language Processing, LMU Munich(慕尼黑大学信息与语言处理中心) Munich Center for Machine Learning (MCML), Munich, Germany(慕尼黑机器学习中心) Free University of Bozen-Bolzano, Italy(博兹纳自由大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

Comments EMNLP 2025 Main, 38 pages, 33 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21318 2025-09-26 cs.CV cs.AI 57%

SD3.5-Flash: Distribution-Guided Distillation of Generative Flows

Hmrishav Bandyopadhyay, Rahim Entezari, Jim Scott, Reshinth Adithyan, Yi-Zhe Song, Varun Jampani

机构 * Stability AI SketchX, University of Surrey(SketchX,大学)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

Comments Project Page: https://hmrishavbandy.github.io/sd35flash/

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21310 2025-09-26 cs.AI 57%

SAGE: A Realistic Benchmark for Semantic Understanding

Samarth Goel, Reagan J. Lee, Kannan Ramchandran

机构 * University of California, Berkeley(加州大学伯克利分校)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

Comments 39th Conference on Neural Information Processing Systems (NeurIPS 2025) Workshop: Evaluating the Evolving LLM Lifecycle: Benchmarks, Emergent Abilities, and Scaling

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21208 2025-09-26 cs.CL 57%

CLaw: Benchmarking Chinese Legal Knowledge in Large Language Models - A Fine-grained Corpus and Reasoning Analysis

Xinzhe Xu, Liang Zhao, Hongshen Xu, Chen Chen

机构 * Peking University(北京大学) LLM-Core

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.20418 2025-09-26 cs.CR cs.AI cs.ET 57%

A Taxonomy of Data Risks in AI and Quantum Computing (QAI) - A Systematic Review

Grace Billiris, Asif Gill, Madhushi Bandara

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

Comments 11 pages, 2 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19096 2025-09-26 cs.CV cs.SE 50%

Investigating Traffic Accident Detection Using Multimodal Large Language Models

Ilhan Skender, Kailin Tong, Selim Solmaz, Daniel Watzenig

机构 * Embedded Systems Group (Dept.-E)(嵌入式系统组) Virtual Vehicle Research GmbH(虚拟车辆研究公司) Control Systems Group (Dept.-E)(控制系统组) Institute of Visual Computing(视觉计算研究所) Graz University of Technology(格拉茨技术大学)

专题命中 安全评测 :safety(abstract)

Comments Accepted for presentation at the 2025 IEEE International Automated Vehicle Validation Conference (IAVVC 2025). Final version to appear in IEEE Xplore

详情

展开后加载摘要…

URL PDF HTML 收藏

7. AI治理与伦理 3 篇

2509.20394 2025-09-26 cs.CY cs.AI cs.CL cs.CR 75%

Blueprints of Trust: AI System Cards for End to End Transparency and Governance

Huzaifa Sidhpurwala, Emily Fox, Garth Mollett, Florencio Cano Gabarda, Roman Zhukov

专题命中 AI治理与伦理 :safety(abstract);AI safety(abstract);分类 cs.CL、cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏