arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-09-05 至 2025-09-05 共收录 29 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全训练 4 篇

2509.04413 2025-09-05 eess.SY cs.LG cs.MA cs.RO cs.SY math.OC 79%

SAFE--MA--RRT: Multi-Agent Motion Planning with Data-Driven Safety Certificates

Babak Esmaeili, Hamidreza Modares

机构 * Department of Mechanical Engineering, Michigan State University(机械工程系,密歇根州立大学)

专题命中 安全训练 :safety(title,abstract);分类 cs.LG

Comments Submitted to IEEE Transactions on Automation Science and Engineering

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.14701 2025-09-05 cs.CL cs.LG 62%

An Unsupervised Natural Language Processing Pipeline for Assessing Referral Appropriateness

Vittorio Torri, Annamaria Bottelli, Michele Ercolanoni, Olivia Leoni, Francesca Ieva

机构 * ARIA s.p.a - Azienda Regionale per l’Innovazione e gli Acquisti(区域创新与采购公司) Regione Lombardia(伦巴第大区) Human Technopole(人类技术极地)

专题命中 安全训练 :alignment(abstract);分类 cs.CL、cs.LG

Comments 49 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.03768 2025-09-05 cs.AI stat.ML 57%

RAGuard: A Novel Approach for in-context Safe Retrieval Augmented Generation for LLMs

Connor Walker, Koorosh Aslansefat, Mohammad Naveed Akram, Yiannis Papadopoulos

机构 * University of Hull(赫尔大学) AURA CDT(AURA联合培养计划)

专题命中 安全训练 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.03550 2025-09-05 cs.AI 57%

Diffusion-RL Based Air Traffic Conflict Detection and Resolution Method

Tonghe Li, Jixin Liu, Weili Zeng, Hao Jiang

专题命中 安全训练 :safety(abstract);分类 cs.AI

Comments 59 pages,13 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 越狱攻击 1 篇

2508.20038 2025-09-05 cs.CL 89%

Forewarned is Forearmed: Pre-Synthesizing Jailbreak-like Instructions to Enhance LLM Safety Guardrail to Potential Attacks

Sheng Liu, Qiang Sheng, Danding Wang, Yang Li, Guang Yang, Juan Cao

机构 * Sheng Liu Media Synthesis and Forensics Lab, Institute of Computing Technology, Chinese Academy of Sciences University of Chinese Academy of Sciences(媒体合成与取证实验室,计算技术研究所,中国科学院,中国科学院大学) Qiang Sheng Media Synthesis and Forensics Lab, Institute of Computing Technology, Chinese Academy of Sciences(媒体合成与取证实验室,计算技术研究所,中国科学院) Danding Wang Media Synthesis and Forensics Lab, Institute of Computing Technology, Chinese Academy of Sciences(媒体合成与取证实验室,计算技术研究所,中国科学院) Yang Li Media Synthesis and Forensics Lab, Institute of Computing Technology, Chinese Academy of Sciences University of Chinese Academy of Sciences(媒体合成与取证实验室,计算技术研究所,中国科学院,中国科学院大学) Guang Yang Zhongguancun Laboratory(中关村实验室) Juan Cao Media Synthesis and Forensics Lab, Institute of Computing Technology, Chinese Academy of Sciences(媒体合成与取证实验室,计算技术研究所,中国科学院)

专题命中 越狱攻击 :safety(title,abstract);jailbreak(title,abstract);alignment(abstract);分类 cs.CL

Comments EMNLP 2025 findings

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 幻觉与事实性 3 篇

2509.03871 2025-09-05 cs.CL cs.AI cs.CR 79%

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models

Yanbo Wang, Yongcan Yu, Jian Liang, Ran He

机构 * School of Artificial Intelligence, University of Chinese Academy of Sciences(人工智能学院,中国科学院大学) NLPR & MAIS, Institute of Automation, Chinese Academy of Sciences(神经语言处理与机器智能中心,中国科学院自动化研究所)

专题命中 幻觉与事实性 :safety(abstract);trustworthy(abstract);AI safety(abstract);分类 cs.CL、cs.AI

Comments 38 pages. This survey considers papers published up to June 30, 2025. Work in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04052 2025-09-05 cs.IR 50%

Safeguarding Patient Trust in the Age of AI: Tackling Health Misinformation with Explainable AI

Sueun Hong, Shuojie Fu, Ovidiu Serban, Brianna Bao, James Kinross, Francesa Toni, Guy Martin, Uddhav Vaghela

专题命中 幻觉与事实性 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.03693 2025-09-05 cs.HC cs.MM 50%

Designing Effective AI Explanations for Misinformation Detection: A Comparative Study of Content, Social, and Combined Explanations

Yeaeun Gong, Yifan Liu, Lanyu Shang, Na Wei, Dong Wang

专题命中 幻觉与事实性 :alignment(abstract)

Comments To appear at CSCW 2025

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 隐私与版权 2 篇

2505.14585 2025-09-05 cs.CL 79%

Context Reasoner: Incentivizing Reasoning Capability for Contextualized Privacy and Safety Compliance via Reinforcement Learning

Wenbin Hu, Haoran Li, Huihao Jing, Qi Hu, Ziqian Zeng, Sirui Han, Heli Xu, Tianshu Chu, Peizhao Hu, Yangqiu Song

机构 * HKUST(香港科技大学) South China University of Technology(华南理工大学) Huawei Technologies(华为技术有限公司)

专题命中 隐私与版权 :safety(title,abstract);分类 cs.CL

Comments Accepted to EMNLP 2025 Main

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00073 2025-09-05 cs.LG 57%

Mitigating Clinician Information Overload: Generative AI for Integrated EHR and RPM Data Analysis

Ankit Shetgaonkar, Dipen Pradhan, Lakshit Arora, Sanjay Surendranath Girija, Shashank Kapoor, Aman Raj

专题命中 隐私与版权 :safety(abstract);分类 cs.LG

Comments Accepted at IEEE COMPSAC 2025

Journal ref 2025 IEEE 49th Annual Computers, Software, and Applications Conference (COMPSAC)

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 安全评测 8 篇

2509.01185 2025-09-05 cs.CL cs.AI 73%

Modular Techniques for Synthetic Long-Context Data Generation in Language Model Training and Evaluation

Seganrasan Subramanian, Abhigya Verma

机构 * ServiceNow

专题命中 安全评测 :alignment(abstract);DPO(abstract);分类 cs.CL、cs.AI

Comments 26 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.03986 2025-09-05 cs.CV cs.AI cs.CL cs.LG 67%

Promptception: How Sensitive Are Large Multimodal Models to Prompts?

Mohamed Insaf Ismithdeen, Muhammad Uzair Khattak, Salman Khan

机构 * Mohamed Bin Zayed University of Artificial Intelligence(莫罕默德·本·扎耶德人工智能大学) Swiss Federal Institute of Technology Lausanne (EPFL)(洛桑联邦理工学院) Australian National University(澳大利亚国立大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments Accepted to EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.03809 2025-09-05 cs.CL cs.AI 62%

Align-then-Slide: A complete evaluation framework for Ultra-Long Document-Level Machine Translation

Jiaxin Guo, Daimeng Wei, Yuanchang Luo, Xiaoyu Chen, Zhanglin Wu, Huan Yang, Hengchao Shang, Zongyao Li, Zhiqiang Rao, Jinlong Yang, Hao Yang

机构 * Huawei Translation Services Center(华为翻译服务中心)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

Comments under preview

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.03529 2025-09-05 cs.CL cs.AI eess.AS 62%

Multimodal Proposal for an AI-Based Tool to Increase Cross-Assessment of Messages

Alejandro Álvarez Castro, Joaquín Ordieres-Meré

机构 * AI master(人工智能硕士) Universidad Politécnica de Madrid(马德里理工大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

Comments Presented at NLMLT2025 (https://airccse.org/csit/V15N16.html), 15 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04292 2025-09-05 cs.CL 57%

Inverse IFEval: Can LLMs Unlearn Stubborn Training Conventions to Follow Real Instructions?

Qinyan Zhang, Xinping Lei, Ruijie Miao, Yu Fu, Haojie Fan, Le Chang, Jiafan Hou, Dingling Zhang, Zhongfei Hou, Ziqiang Yang, Changxin Pu, Fei Hu, Jingkai Liu, Mengyun Liu, Yang Liu, Xiang Gao, Jiaheng Liu, Tong Yang, Zaiyuan Wang, Ge Zhang, Wenhao Huang

机构 * ByteDance Seed(字节跳动种子) Nanjing University(南京大学) Peking University(北京大学) Beijing University of Posts and Telecommunications(北京邮电大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04053 2025-09-05 cs.LG 57%

On Aligning Prediction Models with Clinical Experiential Learning: A Prostate Cancer Case Study

Jacqueline J. Vallon, William Overman, Wanqiao Xu, Neil Panjwani, Xi Ling, Sushmita Vij, Hilary P. Bagshaw, John T. Leppert, Sumit Shah, Geoffrey Sonn, Sandy Srinivas, Erqi Pollom, Mark K. Buyyounouski, Mohsen Bayati

专题命中 安全评测 :alignment(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.03787 2025-09-05 cs.IR cs.CL 57%

Evaluating the Robustness of Retrieval-Augmented Generation to Adversarial Evidence in the Health Domain

Shakiba Amirshahi, Amin Bigdeli, Charles L. A. Clarke, Amira Ghenai

机构 * University of Waterloo(滑铁卢大学) Toronto Metropolitan University(多伦多 Metropolitan 大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.03565 2025-09-05 cs.CL cs.MM 57%

ResearchPulse: Building Method-Experiment Chains through Multi-Document Scientific Inference

Qi Chen, Jingxuan Wei, Zhuoya Yao, Haiguang Wang, Gaowei Wu, Bihui Yu, Siyuan Li, Cheng Tan

机构 * University of Chinese Academy of Sciences(中国科学院大学) Zhejiang University(浙江大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

Comments Accepted to ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏

6. AI治理与伦理 2 篇

2508.08193 2025-09-05 cs.CY cs.AI 62%

Street-Level AI: Are Large Language Models Ready for Real-World Judgments?

Gaurab Pokharel, Shafkat Farabi, Patrick J. Fowler, Sanmay Das

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY

Comments This work has been accepted for publication as a full paper at the AAAI/ACM Conference on AI, Ethics, and Society (AIES 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.02025 2025-09-05 cs.DC cs.AI 57%

Evaluating the Efficacy of LLM-Based Reasoning for Multiobjective HPC Job Scheduling

Prachi Jadhav, Hongwei Jin, Ewa Deelman, Prasanna Balaprakash

机构 * University of Tennessee, Knoxville\ Ridge National Laboratory Oak Ridge, TN USA Argonne National Laboratory Lemont, IL USA University of Southern California Los Angeles, CA USA Oak Ridge National Laboratory Oak Ridge, TN USA University of Tennessee, Knoxville\ Ridge National Laboratory Argonne National Laboratory University of Southern California Oak Ridge National Laboratory

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI

Comments 10 pages, 6 figures, work under review

详情

展开后加载摘要…

URL PDF HTML 收藏

7. 其他安全 9 篇

2509.03934 2025-09-05 cs.CL cs.AI 81%

SelfAug: Mitigating Catastrophic Forgetting in Retrieval-Augmented Generation via Distribution Self-Alignment

Yuqing Huang, Rongyang Zhang, Qimeng Wang, Chengqiang Lu, Yan Gao, Yi Wu, Yao Hu, Xuyang Zhi, Guiquan Liu, Xin Li, Hao Wang, Enhong Chen

机构 * University of Science and Technology of China(中国科学技术大学) Xiaohongshu Inc.(小红书公司)

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.03805 2025-09-05 cs.CL cs.AI 62%

Measuring How (Not Just Whether) VLMs Build Common Ground

Saki Imai, Mert İnan, Anthony Sicilia, Malihe Alikhani

机构 * Northeastern University(东北大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.03680 2025-09-05 cs.GR cs.AI cs.CV 57%

LuxDiT: Lighting Estimation with Video Diffusion Transformer

Ruofan Liang, Kai He, Zan Gojcic, Igor Gilitschenski, Sanja Fidler, Nandita Vijaykumar, Zian Wang

机构 * NVIDIA University of Toronto(多伦多大学) Vector Institute(向量研究所)

专题命中 其他安全 :alignment(abstract);分类 cs.AI

Comments Project page: https://research.nvidia.com/labs/toronto-ai/LuxDiT/

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.08537 2025-09-05 cs.LG math.CT 57%

Recursive Reward Aggregation

Yuting Tang, Yivan Zhang, Johannes Ackermann, Yu-Jie Zhang, Soichiro Nishimori, Masashi Sugiyama

专题命中 其他安全 :alignment(abstract);分类 cs.LG

Comments Reinforcement Learning Conference 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.10118 2025-09-05 cs.CV cs.AI 57%

Image Embedding Sampling Method for Diverse Captioning

Sania Waheed, Na Min An

机构 * University of Southhampton(南安普顿大学) KAIST(韩国科学技术院)

专题命中 其他安全 :alignment(abstract);分类 cs.AI

Comments 17 pages, 5 figures, 9 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04324 2025-09-05 cs.RO cs.CV 50%

OVGrasp: Open-Vocabulary Grasping Assistance via Multimodal Intent Detection

Chen Hu, Shan Luo, Letizia Gionfrida

机构 * Department of Informatics, King's College London(伦敦国王学院信息学院) Department of Engineering, King's College London(伦敦国王学院工程学院) John A. Paulson School of Engineering and Applied Sciences, Harvard University(哈佛大学约翰·A·保罗森工程与应用科学学院)

专题命中 其他安全 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04173 2025-09-05 cs.AR 50%

Real-time Object Detection and Associated Hardware Accelerators Targeting Autonomous Vehicles: A Review

Safa Sali, Anis Meribout, Ashiyana Majeed, Mahmoud Meribout, Juan Pablo, Varun Tiwari, Asma Baobaid

专题命中 其他安全 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00300 2025-09-05 cs.CR 50%

ShadowScope: GPU Monitoring and Validation via Composable Side Channel Signals

Ghadeer Almusaddar, Yicheng Zhang, Saber Ganjisaffar, Barry Williams, Yu David Liu, Dmitry Ponomarev, Nael Abu-Ghazaleh

专题命中 其他安全 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17832 2025-09-05 cs.CV 50%

HLG: Comprehensive 3D Room Construction via Hierarchical Layout Generation

Xiping Wang, Yuxi Wang, Mengqi Zhou, Junsong Fan, Zhaoxiang Zhang

专题命中 其他安全 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏