arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-07-31 至 2025-07-31 共收录 24 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 偏好对齐 3 篇

2507.21391 2025-07-31 cs.CV cs.AI cs.CL 73%

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation

Shijie Zhou, Ruiyi Zhang, Huaisheng Zhu, Branislav Kveton, Yufan Zhou, Jiuxiang Gu, Jian Chen, Changyou Chen

机构 * University at Buffalo(布法罗大学) Adobe Research(Adobe研究) Pennsylvania State University(宾夕法尼亚州立大学)

专题命中 偏好对齐 :alignment(abstract);safety(abstract);分类 cs.CL、cs.AI

Comments Accepted at ICCV 2025. Code available at https://github.com/sjz5202/LLaVA-Reward

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.04858 2025-07-31 cs.AI cs.LG 62%

Don't Lag, RAG: Training-Free Adversarial Detection Using RAG

Roie Kazoom, Raz Lapid, Moshe Sipper, Ofer Hadar

机构 * Electrical and Computer Engineering, Ben Gurion University, Beer Sheba 84105, Israel(电子与计算机工程系,本· Gurion 大学) Computer Science, Ben Gurion University, Beer Sheba 84105, Israel(计算机科学系,本· Gurion 大学)

专题命中 偏好对齐 :DPO(abstract);分类 cs.AI、cs.LG

Comments Accepted at VecDB @ ICML 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.09213 2025-07-31 cs.CL 57%

FineMedLM-o1: Enhancing Medical Knowledge Reasoning Ability of LLM from Supervised Fine-Tuning to Test-Time Training

Hongzhou Yu, Tianhao Cheng, Yingwen Wang, Wen He, Qing Wang, Ying Cheng, Yuejie Zhang, Rui Feng, Xiaobo Zhang

机构 * Fudan University(复旦大学) Children’s Hospital of Fudan University(复旦大学附属儿童医院)

专题命中 偏好对齐 :DPO(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 安全训练 1 篇

2507.22358 2025-07-31 cs.AI cs.HC 57%

Magentic-UI: Towards Human-in-the-loop Agentic Systems

Hussein Mozannar, Gagan Bansal, Cheng Tan, Adam Fourney, Victor Dibia, Jingya Chen, Jack Gerrits, Tyler Payne, Matheus Kunzler Maldaner, Madeleine Grunde-McLaughlin, Eric Zhu, Griffin Bassman, Jacob Alber, Peter Chang, Ricky Loynd, Friederike Niedtner, Ece Kamar, Maya Murad, Rafah Hosn, Saleema Amershi

机构 * Microsoft Research AI Frontiers(微软研究院人工智能前沿)

专题命中 安全训练 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 越狱攻击 1 篇

2507.22304 2025-07-31 cs.CR 50%

Invisible Injections: Exploiting Vision-Language Models Through Steganographic Prompt Embedding

Chetan Pathade

专题命中 越狱攻击 :prompt injection(abstract)

Comments 14 Pages

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 红队测试 1 篇

2507.22133 2025-07-31 cs.CR cs.CL 79%

Prompt Optimization and Evaluation for LLM Automated Red Teaming

Michael Freenor, Lauren Alvarez, Milton Leal, Lily Smith, Joel Garrett, Yelyzaveta Husieva, Madeline Woodruff, Ryan Miller, Erich Kummerfeld, Rafael Medeiros, Sander Schulhoff

机构 * Fuel iX Applied Research(Fuel iX应用研究) North Carolina State University(北卡罗来纳州立大学) University of Minnesota(明尼苏达大学) TELUS Digital(TELUS数字) Learn Prompting

专题命中 红队测试 :red teaming(title,abstract);分类 cs.CL

Comments 9 pages, 5 Figures, and 1 Appendix item

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 幻觉与事实性 1 篇

2507.22282 2025-07-31 cs.MA cs.RO 71%

Multi-Agent Path Finding Among Dynamic Uncontrollable Agents with Statistical Safety Guarantees

Kegan J. Strawn, Thomy Phan, Eric Wang, Nora Ayanian, Sven Koenig, Lars Lindemann

专题命中 幻觉与事实性 :safety(title)

Comments 9 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏

6. 安全评测 8 篇

2403.16591 2025-07-31 cs.LG cs.AI cs.CR 81%

Bridging Privacy and Robustness for Trustworthy Machine Learning

Xiaojin Zhang, Wei Chen

机构 * Huazhong University of Science and Technology(华中科技大学)

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.22389 2025-07-31 cs.RO cs.SY eess.SY 78%

Safety Evaluation of Motion Plans Using Trajectory Predictors as Forward Reachable Set Estimators

Kaustav Chakraborty, Zeyuan Feng, Sushant Veer, Apoorva Sharma, Wenhao Ding, Sever Topan, Boris Ivanovic, Marco Pavone, Somil Bansal

机构 * Department of Electrical Engineering, University of Southern California(电气工程系,美国南加州大学) Department of Aeronautics and Astronautics, Stanford University(航空与宇航系,斯坦福大学) NVIDIA Research(NVIDIA研究)

专题命中 安全评测 :safety(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21919 2025-07-31 cs.CL cs.AI cs.CY 67%

Training language models to be warm and empathetic makes them less reliable and more sycophantic

Lujain Ibrahim, Franziska Sofia Hafner, Luc Rocher

机构 * Oxford Internet Institute(牛津互联网研究所) University of Oxford(牛津大学)

专题命中 安全评测 :safety(abstract);分类 cs.CL、cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.22576 2025-07-31 cs.CV cs.AI cs.LG 62%

COOkeD: Ensemble-based OOD detection in the era of zero-shot CLIP

Galadrielle Humblot-Renaux, Gianni Franchi, Sergio Escalera, Thomas B. Moeslund

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG

Comments accepted at ICCVW'25 - Systematic Trust in AI Models: Ensuring Fairness, Reliability, Explainability, and Accountability in Machine Learning Frameworks

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.01282 2025-07-31 cs.CL 57%

Prompt-Reverse Inconsistency: LLM Self-Inconsistency Beyond Generative Randomness and Prompt Paraphrasing

Jihyun Janice Ahn, Wenpeng Yin

机构 * Department of Computer Science & Engineering(计算机科学与工程系) The Pennsylvania State University(宾夕法尼亚州立大学)

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL

Comments accepted in COLM2025, 9 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.18044 2025-07-31 cs.LG 57%

The Geometry of Queries: Query-Based Innovations in Retrieval-Augmented Generation for Healthcare QA

Eric Yang, Jonathan Amar, Jong Ha Lee, Bhawesh Kumar, Yugang Jia

机构 * Verily Life Sciences(Verily 生物科技)

专题命中 安全评测 :trustworthy(abstract);分类 cs.LG

Comments 27 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.22076 2025-07-31 cs.LG 57%

Test-time Prompt Refinement for Text-to-Image Models

Mohammad Abdul Hafeez Khan, Yash Jain, Siddhartha Bhattacharyya, Vibhav Vineet

机构 * Florida Institute of Technology(佛罗里达理工学院) Microsoft Research(微软研究院)

专题命中 安全评测 :alignment(abstract);分类 cs.LG

Comments Accepted to ICCV 2025, MARS2 Workshop. Total 14 pages, 12 figures and 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.22100 2025-07-31 cs.CV 50%

Trade-offs in Image Generation: How Do Different Dimensions Interact?

Sicheng Zhang, Binzhu Xie, Zhonghao Yan, Yuli Zhang, Donghao Zhou, Xiaofei Chen, Shi Qiu, Jiaqi Liu, Guoyang Xie, Zhichao Lu

机构 * Khalifa University(卡利法大学) The Chinese University of Hong Kong(香港中文大学) Queen Mary University of London(伦敦大学玛丽女王学院) Xi’an Jiaotong-Liverpool University(西安交通大学利物浦大学) City University of Hong Kong(香港城市大学)

专题命中 安全评测 :alignment(abstract)

Comments Accepted in ICCV 2025, Codebase: https://github.com/fesvhtr/TRIG

详情

展开后加载摘要…

URL PDF HTML 收藏

7. AI治理与伦理 4 篇

2507.22326 2025-07-31 cs.AI 83%

An Explainable Emotion Alignment Framework for LLM-Empowered Agent in Metaverse Service Ecosystem

Qun Ma, Xiao Xue, Ming Zhang, Yifan Shen, Zihan Zhao

机构 * College of Intelligence and Computing(智能与计算学院) Tianjin University(天津大学) Tianjin Key Laboratory of Healhy Habitat and Smart Technology(天津健康人居环境与智能技术重点实验室) Laboratory of Computation and Analytics of Complex Management Systems(复杂管理系统计算与分析实验室) Faculty of Environment, Science and Economy(环境、科学与经济学院)

专题命中 AI治理与伦理 :alignment(title,abstract);safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.22267 2025-07-31 cs.HC cs.AI 79%

Promoting Online Safety by Simulating Unsafe Conversations with LLMs

Owen Hoffman, Kangze Peng, Zehua You, Sajid Kamal, Sukrit Venkatagiri

机构 * Department of Computer Science, Swarthmore College(计算机科学系,斯沃斯里学院)

专题命中 AI治理与伦理 :safety(title,abstract);分类 cs.AI

Journal ref ACM 2025 Conference on Conversational User Interfaces Workshop on Personas Evolved: Designing Ethical LLM-Based Conversational Agent Personalities

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20014 2025-07-31 cs.CR cs.AI 70%

Policy-Driven AI in Dataspaces: Taxonomy, Explainability, and Pathways for Compliant Innovation

Joydeep Chandra, Satyam Kumar Navneet

机构 * Department of CST Tsinghua University Beijing, China(计算机科学与技术系 清华大学 北京中国) Department of CSE Chandigarh University Mohali, India(计算机科学与工程系 印度昌迪加尔大学 摩哈利)

专题命中 AI治理与伦理 :alignment(abstract);trustworthy(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.22445 2025-07-31 cs.CL cs.AI 62%

AI-generated stories favour stability over change: homogeneity and cultural stereotyping in narratives generated by gpt-4o-mini

Jill Walker Rettberg, Hermann Wigers

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI

Comments This project has received funding from the European Union's Horizon 2020 research and innovation programme under grant agreement number 101142306. The project is also supported by the Center for Digital Narrative, which is funded by the Research Council of Norway through its Centres of Excellence scheme, project number 332643

Journal ref Open Research Europe 2025, 5:202 [version 1; peer review: awaiting peer review]

详情

展开后加载摘要…

URL PDF HTML 收藏

8. 其他安全 5 篇

2507.22511 2025-07-31 cs.ET 78%

Green Wave as an Integral Part for the Optimization of Traffic Efficiency and Safety: A Survey

Kranthi Kumar Talluri, Christopher Stang, Galia Weidl

专题命中 其他安全 :safety(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20930 2025-07-31 cs.CL cs.AI cs.LG 67%

FRED: Financial Retrieval-Enhanced Detection and Editing of Hallucinations in Language Models

Likun Tan, Kuan-Wei Huang, Kevin Wu

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.22529 2025-07-31 cs.LG cs.AI 62%

Accident-Driven Congestion Prediction and Simulation: An Explainable Framework Using Advanced Clustering and Bayesian Networks

Kranthi Kumar Talluri, Galia Weidl, Vaishnavi Kasuluru

机构 * Aschaffenburg University of Applied Sciences(阿施法亨堡应用科学大学) Centre Tecnològic de Telecomunicacions de Catalunya (CTTC)(加泰罗尼亚电信技术中心(CTTC))

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.22346 2025-07-31 cs.CV 50%

DeltaVLM: Interactive Remote Sensing Image Change Analysis via Instruction-guided Difference Perception

Pei Deng, Wenqian Zhou, Hanlin Wu

机构 * School of Information Science and Technology, Beijing Foreign Studies University(信息科学与技术学院,北京外国语大学)

专题命中 其他安全 :alignment(abstract)

Comments 12 pages, 5 figures. Submitted to IEEE Transactions on Geoscience and Remote Sensing (TGRS). Code and dataset are available at https://github.com/hanlinwu/DeltaVLM

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.22066 2025-07-31 cs.SE cs.CR 50%

CodableLLM: Automating Decompiled and Source Code Mapping for LLM Dataset Generation

Dylan Manuel, Paul Rad

专题命中 其他安全 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏