arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-07-22 至 2025-07-22 共收录 57 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 偏好对齐 7 篇

2410.10148 2025-07-22 cs.LG cs.AI cs.CL 83%

AlphaDPO: Adaptive Reward Margin for Direct Preference Optimization

Junkang Wu, Xue Wang, Zhengyi Yang, Jiancan Wu, Jinyang Gao, Bolin Ding, Xiang Wang, Xiangnan He

机构 * MoE Key Lab of BIPC, University of Science(BIPC联合实验室,科学与技术大学) Alibaba Group(阿里巴巴集团)

专题命中 偏好对齐 :alignment(abstract);RLHF(abstract);DPO(abstract);safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.17141 2025-07-22 cs.CL cs.AI 81%

MetaAligner: Towards Generalizable Multi-Objective Alignment of Language Models

Kailai Yang, Zhiwei Liu, Qianqian Xie, Jimin Huang, Tianlin Zhang, Sophia Ananiadou

机构 * The University of Manchester(曼彻斯特大学)

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.CL、cs.AI

Comments Accepted by NeurIPS 2024 main track

Journal ref https://proceedings.neurips.cc/paper_files/paper/2024/hash/3d03800841fa1bb2f43ef1750aafcce4-Abstract-Conference.html

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.15507 2025-07-22 cs.LG cs.AI cs.CL 67%

Off-Policy Corrected Reward Modeling for Reinforcement Learning from Human Feedback

Johannes Ackermann, Takashi Ishida, Masashi Sugiyama

专题命中 偏好对齐 :RLHF(abstract);分类 cs.CL、cs.AI、cs.LG

Comments Accept at the Conference On Language Modeling (COLM) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.15281 2025-07-22 cs.CL cs.AI 62%

A Novel Self-Evolution Framework for Large Language Models

Haoran Sun, Zekun Zhang, Shaoning Zeng

机构 * Haoran Sun, Zekun Zhang, Shaoning Zeng(作者)

专题命中 偏好对齐 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.15275 2025-07-22 cs.CL 57%

ChiMed 2.0: Advancing Chinese Medical Dataset in Facilitating Large Language Modeling

Yuanhe Tian, Junjie Liu, Zhizhou Kou, Yuxiang Li, Yan Song

机构 * University of Washington(华盛顿大学) University of Science and Technology of China(中国科学技术大学)

专题命中 偏好对齐 :RLHF(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.02065 2025-07-22 cs.CR cs.AI cs.IR 57%

Too Much to Trust? Measuring the Security and Cognitive Impacts of Explainability in AI-Driven SOCs

Nidhi Rastogi, Shirid Pant, Devang Dhanuka, Amulya Saxena, Pranjal Mairal

机构 * Rochester Institute of Technology(罗切斯特技术学院) Independent Researcher(独立研究者)

专题命中 偏好对齐 :trustworthy(abstract);分类 cs.AI

Comments 13 pages, ACM Conference on Computer and Communications Security (CCS), 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.04695 2025-07-22 cs.LG 57%

Interpretable Reward Modeling with Active Concept Bottlenecks

Sonia Laguna, Katarzyna Kobalczyk, Julia E. Vogt, Mihaela Van der Schaar

机构 * Department of Computer Science, ETH Zurich, Zurich, Switzerland(苏黎世联邦理工学院计算机科学系) Department of Applied Mathematics and Theoretical Physics, University of Cambridge, United Kingdom(剑桥大学应用数学与理论物理系)

专题命中 偏好对齐 :RLHF(abstract);分类 cs.LG

Journal ref ICML 2025 Workshop on Programmatic Representations for Agent Learning

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 安全训练 5 篇

2507.14987 2025-07-22 cs.AI cs.CR cs.LG 88%

AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning

Yi Zhang, An Zhang, XiuYu Zhang, Leheng Sheng, Yuxin Chen, Zhenkai Liang, Xiang Wang

专题命中 安全训练 :alignment(title,abstract);safety(title,abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14207 2025-07-22 cs.CR cs.AI 70%

Mitigating Trojanized Prompt Chains in Educational LLM Use Cases: Experimental Findings and Detection Tool Design

Richard M. Charles, James H. Curry, Richard B. Charles

机构 * University of Colorado, Boulder(科罗拉多大学波德罗尔分校) Charles Analytics, Aurora(查尔斯分析公司)

专题命中 安全训练 :safety(abstract);AI safety(abstract);分类 cs.AI

Comments 12 pages, 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14293 2025-07-22 cs.AI cs.CL cs.CV 62%

WebGuard: Building a Generalizable Guardrail for Web Agents

Boyuan Zheng, Zeyi Liao, Scott Salisbury, Zeyuan Liu, Michael Lin, Qinyuan Zheng, Zifan Wang, Xiang Deng, Dawn Song, Huan Sun, Yu Su

机构 * The Ohio State University(俄亥俄州立大学) Scale AI University of California, Berkeley(加州大学伯克利分校)

专题命中 安全训练 :safety(abstract);分类 cs.CL、cs.AI

Comments We publicly release WebGuard, along with its annotation tools and fine-tuned models, to facilitate open-source research on monitoring and safeguarding web agents. All resources are available at https://github.com/OSU-NLP-Group/WebGuard

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14438 2025-07-22 physics.ins-det hep-ex 50%

A GEANT4-Based Simulation of Directional Neutron Detectors Using Liquid Scintillators and Boron Carbide Moderators

J. -H. Chen, M. Mirzakhani, R. Mahapatra, S. Sahoo

专题命中 安全训练 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.17646 2025-07-22 cs.PL 50%

Portability of Optimizations from SC to TSO

Akshay Gopalakrishnan, Clark Verbrugge

专题命中 安全训练 :safety(abstract)

Comments Submitted Manuscript. This pre-print has not undergone any post-review modifications/improvements

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 红队测试 2 篇

2507.14202 2025-07-22 cs.CR cs.AI 88%

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training

Pengfei Du

机构 * Pengfei Du(独立研究者)

专题命中 红队测试 :alignment(title,abstract);red teaming(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.15587 2025-07-22 cs.LG cs.AI 62%

Red-Team Multi-Agent Reinforcement Learning for Emergency Braking Scenario

Yinsong Chen, Kaifeng Wang, Xiaoqiang Meng, Xueyuan Li, Zirui Li, Xin Gao

机构 * School of Mechanical Engineering Beijing Institute of Technology Beijing, China(机械工程学院 北京理工大学 北京中国)

专题命中 红队测试 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 提示注入 5 篇

2507.15219 2025-07-22 cs.CR cs.AI 79%

PromptArmor: Simple yet Effective Prompt Injection Defenses

Tianneng Shi, Kaijie Zhu, Zhun Wang, Yuqi Jia, Will Cai, Weida Liang, Haonan Wang, Hend Alzahrani, Joshua Lu, Kenji Kawaguchi, Basel Alomair, Xuandong Zhao, William Yang Wang, Neil Gong, Wenbo Guo, Dawn Song

机构 * UC Berkeley(加州大学伯克利分校) UC Santa Barbara(加州大学圣巴巴拉分校) Duke University(杜克大学) National University of Singapore(新加坡国立大学) King Abdulaziz City for Science and Technology(国王阿卜杜勒阿齐兹城市科学技术学院) University of Washington(华盛顿大学)

专题命中 提示注入 :prompt injection(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14799 2025-07-22 cs.CR cs.AI 79%

Manipulating LLM Web Agents with Indirect Prompt Injection Attack via HTML Accessibility Tree

Sam Johnson, Viet Pham, Thai Le

专题命中 提示注入 :prompt injection(title,abstract);分类 cs.AI

Comments EMNLP 2025 System Demonstrations Submission

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.15613 2025-07-22 cs.CR cs.AI 70%

Multi-Stage Prompt Inference Attacks on Enterprise LLM Systems

Andrii Balashov, Olena Ponomarova, Xiaohua Zhai

机构 * Ukrainian State University of Science and Technologies, ESI "Prydniprovska State Academy of Civil Engineering and Architecture(乌克兰科学与技术国家大学,埃斯I普里皮亚特斯卡州立建筑与工程学院) Google DeepMind(谷歌DeepMind)

专题命中 提示注入 :safety(abstract);prompt injection(abstract);分类 cs.AI

Comments 26 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.15330 2025-07-22 cs.AI 57%

QSAF: A Novel Mitigation Framework for Cognitive Degradation in Agentic AI

Hammad Atta, Muhammad Zeeshan Baig, Yasir Mehmood, Nadeem Shahzad, Ken Huang, Muhammad Aziz Ul Haq, Muhammad Awais, Kamal Ahmed

机构 * Qorvex Consulting(Qorvex咨询公司) Wentworth Institute of Higher Education(文森特高等教育学院) RAN Verification NOKIA, Germany(NOKIA德国RAN验证) Roshan Consulting(罗尚咨询) Skylink Antenna(Skylink天线) Eviden Saudi Arabia(沙特Eviden) Deloitte Enterprise Risk | Internal Audit | Technology GRC(德勤企业风险管理|内部审计|技术治理)

专题命中 提示注入 :prompt injection(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14860 2025-07-22 cs.CY physics.ed-ph 57%

Strategic Integration of AI Chatbots in Physics Teacher Preparation: A TPACK-SWOT Analysis of Pedagogical, Epistemic, and Cybersecurity Dimensions

N. Mohammadipour

专题命中 提示注入 :prompt injection(abstract);分类 cs.CY

Comments 34 pages, 3 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 幻觉与事实性 4 篇

2507.14239 2025-07-22 cs.CL cs.AI 62%

CCL-XCoT: An Efficient Cross-Lingual Knowledge Transfer Method for Mitigating Hallucination Generation

Weihua Zheng, Roy Ka-Wei Lee, Zhengyuan Liu, Kui Wu, AiTi Aw, Bowei Zou

机构 * Institute for Infocomm Research (I 2 R), A*STAR, Singapore(信息与通信研究机构(I2R),A*STAR,新加坡)

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.16215 2025-07-22 cs.AI cs.LG eess.SP 62%

Smarter Together: Combining Large Language Models and Small Models for Physiological Signals Visual Inspection

Huayu Li, Zhengxiao He, Xiwen Chen, Ci Zhang, Stuart F. Quan, William D. S. Killgore, Shu-Fen Wung, Chen X. Chen, Geng Yuan, Jin Lu, Ao Li

机构 * Department of Electrical and Computer Engineering, University of Arizona(电气与计算机工程系,亚利桑那大学) School of Computing, Clemson University(计算学院,克莱姆斯大学) School of Computing, University of Georgia(计算学院,佐治亚大学) Department of Medicine, University of Arizona(医学系,亚利桑那大学) Harvard Medical School and Brigham and Women’s Hospital(哈佛医学院和布里洛妇女医院) Department of Psychiatry, University of Arizona(精神病学系,亚利桑那大学) BIO5 Institute, University of Arizona(BIO5研究所,亚利桑那大学) College of Nursing, University of Arizona(护理学院,亚利桑那大学)

专题命中 幻觉与事实性 :trustworthy(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.19073 2025-07-22 cs.CL 57%

Towards Harmonized Uncertainty Estimation for Large Language Models

Rui Li, Jing Long, Muge Qi, Heming Xia, Lei Sha, Peiyi Wang, Zhifang Sui

机构 * Peking University(北京大学) The Hong Kong Polytechnic University(香港理工大学) Beihang University(北航)

专题命中 幻觉与事实性 :trustworthy(abstract);分类 cs.CL

Comments ACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14782 2025-07-22 stat.ML cs.LG math-ph math.MP stat.CO 57%

Uncertainty Quantification for Machine Learning-Based Prediction: A Polynomial Chaos Expansion Approach for Joint Model and Input Uncertainty Propagation

Xiaoping Du

机构 * School of Mechanical Engineering(机械工程学院) Purdue University(普渡大学)

专题命中 幻觉与事实性 :trustworthy(abstract);分类 cs.LG

Comments This manuscript has been submitted to Multidisciplinary and Structural Optimization

详情

展开后加载摘要…

URL PDF HTML 收藏

6. 隐私与版权 1 篇

2507.14633 2025-07-22 cs.NI cs.LG 57%

Agentic Satellite-Augmented Low-Altitude Economy and Terrestrial Networks: A Survey on Generative Approaches

Xiaozheng Gao, Yichen Wang, Bosen Liu, Xiao Zhou, Ruichen Zhang, Jiacheng Wang, Dusit Niyato, Dong In Kim, Abbas Jamalipour, Chau Yuen, Jianping An, Kai Yang

机构 * School of Information and Electronics, Beijing Institute of Technology(信息与电子学院,北京理工大学) College of Computing and Data Science, Nanyang Technological University(计算与数据科学学院,南洋理工大学) Department of Electrical and Computer Engineering, Sungkyunkwan University(电气与计算机工程系,首尔大学) School of Electrical and Computer Engineering, University of Sydney(电气与计算机工程学院,悉尼大学) School of Electrical and Electronics Engineering, Nanyang Technological University(电气与电子工程学院,南洋理工大学) School of Cyberspace Science and Technology, Beijing Institute of Technology(网络空间科学与技术学院,北京理工大学)

专题命中 隐私与版权 :trustworthy(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏

7. 安全评测 18 篇

2507.14693 2025-07-22 cs.CL cs.AI cs.CY cs.LG 79%

Rethinking Suicidal Ideation Detection: A Trustworthy Annotation Framework and Cross-Lingual Model Evaluation

Amina Dzafic, Merve Kavut, Ulya Bayram

专题命中 安全评测 :trustworthy(title);分类 cs.CL、cs.AI、cs.CY

Comments This manuscript has been submitted to the IEEE Journal of Biomedical and Health Informatics

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14615 2025-07-22 cs.CL cs.AI 73%

Retrieval-Augmented Clinical Benchmarking for Contextual Model Testing in Kenyan Primary Care: A Methodology Paper

Fred Mutisya, Shikoh Gitau, Christine Syovata, Diana Oigara, Ibrahim Matende, Muna Aden, Munira Ali, Ryan Nyotu, Diana Marion, Job Nyangena, Nasubo Ongoma, Keith Mbae, Elizabeth Wamicha, Eric Mibuari, Jean Philbert Nsengemana, Talkmore Chidede

专题命中 安全评测 :alignment(abstract);safety(abstract);分类 cs.CL、cs.AI

Comments 29 pages, 6 figs, 6 tables. Companion methods paper forthcoming

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.15255 2025-07-22 eess.SP cs.AI cs.LG 62%

MEETI: A Multimodal ECG Dataset from MIMIC-IV-ECG with Signals, Images, Features and Interpretations

Deyun Zhang, Xiang Lan, Shijia Geng, Qinghao Zhao, Sumei Fan, Mengling Feng, Shenda Hong

机构 * HeartVoice Medical Technology(HeartVoice医疗科技) Saw Swee Hock School of Public Health and Institute of Data Science(Saw Swee Hock公共卫生学院和数据科学研究所) National University of Singapore(新加坡国立大学) Department of Cardiology, Peking University People’s Hospital(北京大学人民医院心内科) College of Integrative Chinese and Western Medicine, Anhui University of Chinese Medicine(安徽中医药大学整合中西医学学院) National Institute of Health Data Science, Peking University(北京大学国家健康数据科学研究院) Institute for Artificial Intelligence, Peking University(北京大学人工智能研究院)

专题命中 安全评测 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14824 2025-07-22 cs.LG cs.AI 62%

Benchmarking Foundation Models with Multimodal Public Electronic Health Records

Kunyu Yu, Rui Yang, Jingchi Liao, Siqi Li, Huitao Li, Irene Li, Yifan Peng, Rishikesan Kamaleswaran, Nan Liu

机构 * Centre for Quantitative Medicine and Duke-NUS AI + Medical Science Initiative, Duke-NUS Medical School(定量医学中心和杜克-国立新加坡大学AI+医学科学计划,杜克-国立新加坡大学医学院) Graduate School of Engineering, The University of Tokyo(东京大学工程研究生院) Department of Population Health Sciences, Weill Cornell Medicine(流行病学与公共卫生科学系,韦尔·科恩医学中心) Department of Surgery, Duke University School of Medicine(外科医学系,杜克大学医学学院) Centre for Quantitative Medicine, Duke-NUS AI + Medical Science Initiative and Programme in Health Services and Systems Research, Duke-NUS Medical School and NUS Artificial Intelligence Institute, National University of Singapore(定量医学中心和杜克-国立新加坡大学AI+医学科学计划及健康服务与系统研究计划,杜克-国立新加坡大学医学院和新加坡国立大学人工智能研究所)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14298 2025-07-22 cs.CL cs.AI cs.CV 62%

In-Depth and In-Breadth: Pre-training Multimodal Language Models Customized for Comprehensive Chart Understanding

Wan-Cyuan Fan, Yen-Chun Chen, Mengchen Liu, Alexander Jacobson, Lu Yuan, Leonid Sigal

机构 * UBC(不列颠哥伦比亚大学) Microsoft(微软) Vector Institute for AI(人工智能向量研究所) CIFAR AI Chair(卡尔·弗雷德里克人工智能主席)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

Comments arXiv admin note: substantial text overlap with arXiv:2407.14506

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14180 2025-07-22 cs.LG cs.AI 62%

Digital Twin-Assisted Explainable AI for Robust Beam Prediction in mmWave MIMO Systems

Nasir Khan, Asmaa Abdallah, Abdulkadir Celik, Ahmed M. Eltawil, Sinem Coleri

机构 * department of Electrical and Electronics Engineering, Koc University(电子与电气工程系,科克大学) Computer, Electrical, and Mathematical Sciences and Engineering Division, King Abdullah University of Science and Technology(计算机、电气和数学科学与工程系,国王阿卜杜勒-阿齐兹大学)

专题命中 安全评测 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏