arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-08-27 至 2025-08-27 共收录 35 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 偏好对齐 3 篇

2508.18391 2025-08-27 cs.AI 85%

PKG-DPO: Optimizing Domain-Specific AI systems with Physics Knowledge Graphs and Direct Preference Optimization

Nitin Nagesh Kulkarni, Bryson Wilcox, Max Sawa, Jason Thom

机构 * Advanced Engineering and Technology(先进工程与技术)

专题命中 偏好对齐 :DPO(title,abstract);alignment(abstract);safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.10717 2025-08-27 cs.CL cs.AI 81%

A Modular Approach for Clinical SLMs Driven by Synthetic Data with Pre-Instruction Tuning, Model Merging, and Clinical-Tasks Alignment

Jean-Philippe Corbeil, Amin Dada, Jean-Michel Attendu, Asma Ben Abacha, Alessandro Sordoni, Lucas Caccia, François Beaulieu, Thomas Lin, Jens Kleesiek, Paul Vozila

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.CL、cs.AI

Journal ref ACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.18569 2025-08-27 cs.CL cs.CV 57%

The Mind's Eye: A Multi-Faceted Reward Framework for Guiding Visual Metaphor Generation

Girish A. Koushik, Fatemeh Nazarieh, Katherine Birch, Shenbin Qian, Diptesh Kanojia

机构 * NICE Research Group(NICE研究组) Centre for Translation Studies(翻译研究中心) Institute for People-Centred AI(以人为本的人工智能研究所) School of Computer Science & Electronic Engineering, University of Surrey, UK(Surrey大学计算机科学与电子工程学院)

专题命中 偏好对齐 :alignment(abstract);分类 cs.CL

Comments Under Review

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 安全训练 4 篇

2508.18884 2025-08-27 cs.LG cs.AI 62%

HAEPO: History-Aggregated Exploratory Policy Optimization

Gaurish Trivedi, Alakh Sharma, Kartikey Singh Bhandari, Dhruv Kumar, Pratik Narang, Jagat Sesh Challa

机构 * Birla Institute of Technology and Science, Pilani(比拉理工学院和科学学院,比兰)

专题命中 安全训练 :DPO(abstract);分类 cs.AI、cs.LG

Comments Under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.18180 2025-08-27 cs.AI cs.LG 62%

Safe Reinforcement Learning in Black-Box Environments via Adaptive Shielding

Daniel Bethell, Simos Gerasimou, Radu Calinescu, Calum Imrie

机构 * University of York, UK(英国约克大学) Cyprus University of Technology, Cyprus(塞浦路斯技术大学)

专题命中 安全训练 :safety(abstract);分类 cs.AI、cs.LG

Comments To be published in ECAI 25

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.20361 2025-08-27 cs.MA cs.AI cs.RO 57%

Safe Multiagent Coordination via Entropic Exploration

Ayhan Alp Aydeniz, Enrico Marchesini, Robert Loftin, Christopher Amato, Kagan Tumer

专题命中 安全训练 :safety(abstract);分类 cs.AI

Comments 10 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.07300 2025-08-27 cs.CY 57%

From Principles to Rules: A Regulatory Approach for Frontier AI

Jonas Schuett, Markus Anderljung, Alexis Carlier, Leonie Koessler, Ben Garfinkel

专题命中 安全训练 :safety(abstract);分类 cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 越狱攻击 2 篇

2504.21038 2025-08-27 cs.CR cs.AI 85%

Prefill-level Jailbreak: A Black-Box Risk Analysis of Large Language Models

Yakai Li, Jiekang Hu, Weiduan Sang, Luping Ma, Dongsheng Nie, Weijuan Zhang, Aimin Yu, Yi Su, Qingjia Huang, Qihang Zhou

机构 * Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所) School of Cyber Security, University of Chinese Academy of Sciences(中国科学院大学网络安全学院) Sinochem Energy High-Tech Co., Ltd.(中石化能源高科技有限公司)

专题命中 越狱攻击 :jailbreak(title,abstract);alignment(abstract);safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15648 2025-08-27 cs.CL 83%

SDGO: Self-Discrimination-Guided Optimization for Consistent Safety in Large Language Models

Peng Ding, Wen Sun, Dailin Li, Wei Zou, Jiaming Wang, Jiajun Chen, Shujian Huang

专题命中 越狱攻击 :safety(title,abstract);alignment(abstract);分类 cs.CL

Comments Accepted by EMNLP 2025 (Main Conference), 15 pages, 4 figures, 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 幻觉与事实性 1 篇

2508.02319 2025-08-27 cs.CV 67%

Is Uncertainty Quantification a Viable Alternative to Learned Deferral?

Anna M. Wundram, Christian F. Baumgartner

机构 * Faculty of Health Sciences and Medicine, University of Lucerne, Switzerland(健康科学与医学学院,卢塞恩大学,瑞士) Cluster of Excellence -- ML for Science, University of Tübingen, Germany(卓越中心——科学中的机器学习,图宾根大学,德国)

专题命中 幻觉与事实性 :safety(abstract);AI safety(abstract)

Comments Accepted as an oral presentation at MICCAI UNSURE 2025

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 隐私与版权 2 篇

2508.19115 2025-08-27 cs.CR cs.AI 57%

SecureV2X: An Efficient and Privacy-Preserving System for Vehicle-to-Everything (V2X) Applications

Joshua Lee, Ali Arastehfard, Weiran Liu, Xuegang Ban, Yuan Hong

机构 * University of California Santa Barbara(加州大学圣芭芭拉分校) University of Connecticut(康涅狄格大学) Alibaba Group(阿里巴巴集团) University of Washington(华盛顿大学)

专题命中 隐私与版权 :safety(abstract);分类 cs.AI

Comments 10 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11167 2025-08-27 cs.CV 50%

VFM-Guided Semi-Supervised Detection Transformer under Source-Free Constraints for Remote Sensing Object Detection

Jianhong Han, Yupei Wang, Liang Chen

机构 * School of Information and Electronics, Beijing Institute of Technology(信息与电子学院,北京理工大学) Beijing Institute of Technology Chongqing Innovation Center(北京理工大学重庆创新中心) National Key Laboratory for Space-Born Intelligent Information Processing(空间智能信息处理国家重点实验室)

专题命中 隐私与版权 :alignment(abstract)

Comments Manuscript submitted to IEEE TCSVT

详情

展开后加载摘要…

URL PDF HTML 收藏

6. 安全评测 12 篇

2508.19096 2025-08-27 cs.AI 79%

Trustworthy Agents for Electronic Health Records through Confidence Estimation

Yongwoo Song, Minbyul Jeong, Mujeen Sung

机构 * Kyung Hee University(庆熙大学)

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.18891 2025-08-27 cs.LG cs.AI 62%

pyFAST: A Modular PyTorch Framework for Time Series Modeling with Multi-source and Sparse Data

Zhijin Wang, Senzhen Wu, Yue Hu, Xiufeng Liu

机构 * College of Computer Engineering, Jimei University(嘉应大学计算机工程学院) Chengyi College, Jimei University(嘉应大学 Chengyi 学院) Department of Technology, Management and Economics, Technical University of Denmark(丹麦技术大学技术、管理与经济系)

专题命中 安全评测 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.08846 2025-08-27 cs.CR cs.CL cs.LG 62%

Fingerprint Vector: Enabling Scalable and Efficient Model Fingerprint Transfer via Vector Addition

Zhenhua Xu, Qichen Liu, Zhebo Wang, Wenpeng Xing, Dezhang Kong, Mohan Li, Meng Han

专题命中 安全评测 :harmlessness(abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.19163 2025-08-27 cs.AI cs.HC cs.MA 57%

MATRIX: Multi-Agent simulaTion fRamework for safe Interactions and conteXtual clinical conversational evaluation

Ernest Lim, Yajie Vera He, Jared Joselowitz, Kate Preston, Mohita Chowdhury, Louis Williams, Aisling Higham, Katrina Mason, Mariane Melo, Tom Lawton, Yan Jia, Ibrahim Habli

机构 * Ufonia Limited(乌菲尼亚有限公司) University of York(约克大学) NHS Improvement Academy(国家健康服务改进学院)

专题命中 安全评测 :safety(abstract);分类 cs.AI

Comments 36 pages, 16 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.19097 2025-08-27 cs.AI 57%

Reasoning LLMs in the Medical Domain: A Literature Survey

Armin Berger, Sarthak Khanna, David Berghaus, Rafet Sifa

机构 * Fraunhofer IAIS - Department of Media Engineering(弗劳恩霍夫研究所媒体工程部门) University of Bonn - Department of Computer Science(波恩大学计算机科学系) West-AI - Federal Ministry of Education and Research(西德人工智能 - 教育与研究部)

专题命中 安全评测 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.18715 2025-08-27 cs.CL 57%

EMMM, Explain Me My Model! Explainable Machine Generated Text Detection in Dialogues

Angela Yifei Yuan, Haoyi Li, Soyeon Caren Han, Christopher Leckie

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL

Comments 15 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.18687 2025-08-27 cs.CL 57%

Knowing or Guessing? Robust Medical Visual Question Answering via Joint Consistency and Contrastive Learning

Songtao Jiang, Yuxi Chen, Sibo Song, Yan Zhang, Yeying Jin, Yang Feng, Jian Wu, Zuozhu Liu

机构 * Zhejiang University, Zhejiang, China(浙江大学) Alibaba Group, Zhejiang, China(阿里巴巴集团) Angelalign Technology Inc., Shanghai, China(Angelalign技术有限公司) ChohoTech Inc., Hangzhou, China(楚合科技有限公司) Zhejiang Key Laboratory of Medical Imaging Artificial Intelligence, Zhejiang, China(浙江省医学影像人工智能重点实验室)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.18430 2025-08-27 cs.CV cs.AI 57%

CLARIFY: A Specialist-Generalist Framework for Accurate and Lightweight Dermatological Visual Question Answering

Aranya Saha, Tanvir Ahmed Khan, Ismam Nur Swapnil, Mohammad Ariful Haque

机构 * Aranya Saha ∗ , Tanvir Ahmed Khan ∗ , Ismam Nur Swapnil ∗ , and Mohammad Ariful Haque ∗(无明确机构)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

Comments 10 pages, 8 figures, Prepared for submission to IEEE Transactions on Human-Machine Systems

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.18292 2025-08-27 cs.MA cs.AI 57%

Consensus Is All You Need: Gossip-Based Reasoning Among Large Language Models

Saksham Arora

机构 * Sunnyvale, California(加州松林市)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

Comments 4 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.17367 2025-08-27 cs.CV cs.AI 57%

EVM-Fusion: An Explainable Vision Mamba Architecture with Neural Algorithmic Fusion

Zichuan Yang, Yongzhi Wang

机构 * School of Mathematical Sciences, Tongji University(数学科学学院,同济大学)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

Comments 8 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.18725 2025-08-27 cs.NI cs.IT math.IT 50%

Toward Edge General Intelligence with Agentic AI and Agentification: Concepts, Technologies, and Future Directions

Ruichen Zhang, Guangyuan Liu, Yinqiu Liu, Changyuan Zhao, Jiacheng Wang, Yunting Xu, Dusit Niyato, Jiawen Kang, Yonghui Li, Shiwen Mao, Sumei Sun, Xuemin Shen, Dong In Kim

专题命中 安全评测 :trustworthy(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.20353 2025-08-27 cs.RO 50%

Trajectory-to-Action Pipeline (TAP): Automated Scenario Description Extraction for Autonomous Vehicle Behavior Comparison

Aron Harder, Madhur Behl

机构 * Department of Computer Science, University of Virginia(弗吉尼亚大学计算机科学系)

专题命中 安全评测 :safety(abstract)

Comments 11 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏

7. AI治理与伦理 3 篇

2508.18886 2025-08-27 cs.CV 78%

Toward Robust Medical Fairness: Debiased Dual-Modal Alignment via Text-Guided Attribute-Disentangled Prompt Learning for Vision-Language Models

Yuexuan Xia, Benteng Ma, Jiang He, Zhiyong Wang, Qi Dou, Yong Xia

机构 * National Engineering Laboratory for Integrated Aero-Space-Ground-Ocean Big Data Application Technology(集成空天地海大数据应用技术国家工程实验室) Northwestern Polytechnical University(西北工业大学) Huiying Medical Technology Company Ltd.(慧影医疗技术有限公司) The School of Computer Science(计算机学院) The University of Sydney(悉尼大学) Department of Computer Science and Engineering(计算机科学与工程系) The Chinese University of Hong Kong(香港中文大学)

专题命中 AI治理与伦理 :alignment(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.13042 2025-08-27 cs.CY 57%

How Do AI Companies "Fine-Tune" Policy? Examining Regulatory Capture in AI Governance

Kevin Wei, Carson Ezell, Nick Gabrieli, Chinmay Deshpande

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CY

Comments 39 pages (14 pages main text), 3 figures, 9 tables. To be published in the Proceedings of the 2024 AAAI/ACM Conference on AI, Ethics, & Society (AIES)

Journal ref Proc. AAAI/ACM Conf. AI, Ethics & Soc., 7 (2024) 1539-1555

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.18600 2025-08-27 cs.GT cs.MA econ.GN q-fin.EC 50%

Bias-Adjusted LLM Agents for Human-Like Decision-Making via Behavioral Economics

Ayato Kitadai, Yusuke Fukasawa, Nariaki Nishino

专题命中 AI治理与伦理 :alignment(abstract)

Comments 8 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏

8. 其他安全 8 篇

2312.09625 2025-08-27 cs.CV cs.CL 79%

Weakly-Supervised 3D Visual Grounding based on Visual Language Alignment

Xiaoxu Xu, Yitian Yuan, Qiudan Zhang, Wenhui Wu, Zequn Jie, Lin Ma, Xu Wang

机构 * College of Computer Science and Software Engineering, Shenzhen University(深圳大学计算机科学与软件工程学院) Meituan Inc.(美团公司) College of Electronics and Information Engineering, Shenzhen University(深圳大学电子与信息工程学院) Guangdong Key Laboratory of Intelligent Information Processing(广东省智能信息处理重点实验室)

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.16801 2025-08-27 cs.CV 78%

Decoupled Global-Local Alignment for Improving Compositional Understanding

Xiaoxing Hu, Kaicheng Yang, Jun Wang, Haoran Xu, Ziyong Feng, Yupei Wang

机构 * Beijing Institute of Technology(北京理工大学) Zhejiang University(浙江大学)

专题命中 其他安全 :alignment(title,abstract)

Comments ACMMM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.19099 2025-08-27 cs.CL 57%

Beyond the Black Box: Integrating Lexical and Semantic Methods in Quantitative Discourse Analysis with BERTopic

Thomas Compton

机构 * University of York(约克大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL

Comments 5 pages conference paper, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏