arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-07-21 至 2025-07-21 共收录 26 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 偏好对齐 2 篇

2409.04617 2025-07-21 cs.CL 57%

Sparse Rewards Can Self-Train Dialogue Agents

Barrett Martin Lattimer, Varun Gangal, Ryan McDonald, Yi Yang

专题命中 偏好对齐 :alignment(abstract);分类 cs.CL

Comments Accepted to ACL 2025 (Findings)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.13541 2025-07-21 cs.AI 57%

PrefPalette: Personalized Preference Modeling with Latent Attributes

Shuyue Stella Li, Melanie Sclar, Hunter Lang, Ansong Ni, Jacqueline He, Puxin Xu, Andrew Cohen, Chan Young Park, Yulia Tsvetkov, Asli Celikyilmaz

机构 * Meta FAIR University of Washington(华盛顿大学) Meta GenAI

专题命中 偏好对齐 :trustworthy(abstract);分类 cs.AI

Comments 17 pages, 6 tables, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 越狱攻击 2 篇

2507.13474 2025-07-21 cs.CL 83%

Paper Summary Attack: Jailbreaking LLMs through LLM Safety Papers

Liang Lin, Zhihao Xu, Xuehai Tang, Shi Liu, Biyu Zhou, Fuqing Zhu, Jizhong Han, Songlin Hu

机构 * Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所) Renmin University of China(中国人民大学)

专题命中 越狱攻击 :safety(title,abstract);alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.13761 2025-07-21 cs.CL 57%

Innocence in the Crossfire: Roles of Skip Connections in Jailbreaking Visual Language Models

Palash Nandi, Maithili Joshi, Tanmoy Chakraborty

机构 * Department of Electrical Engineering(电气工程系) Indian Institute of Technology Delhi(印度理工学院德里)

专题命中 越狱攻击 :jailbreak(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 幻觉与事实性 1 篇

2507.13411 2025-07-21 cs.CL cs.AI 62%

Aligning Knowledge Graphs and Language Models for Factual Accuracy

Nur A Zarin Nishat, Andrea Coletta, Luigi Bellomarini, Kossi Amouzouvi, Jens Lehmann, Sahar Vahdati

机构 * TIB – Leibniz Information Centre for Science and Technology(莱比锡信息科学与技术研究中心) Amazon, Germany(亚马逊公司)

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 安全评测 15 篇

2507.13383 2025-07-21 cs.LG cs.AI cs.CV 88%

Whose View of Safety? A Deep DIVE Dataset for Pluralistic Alignment of Text-to-Image Models

Charvi Rastogi, Tian Huey Teh, Pushkar Mishra, Roma Patel, Ding Wang, Mark Díaz, Alicia Parrish, Aida Mostafazadeh Davani, Zoe Ashwood, Michela Paganini, Vinodkumar Prabhakaran, Verena Rieser, Lora Aroyo

机构 * Google DeepMind(谷歌DeepMind) Google Research(谷歌研究)

专题命中 安全评测 :alignment(title,abstract);safety(title,abstract);分类 cs.AI、cs.LG

Comments 28 pages, 16 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.02145 2025-07-21 cs.AI cs.CL cs.RO 81%

From Words to Collisions: LLM-Guided Evaluation and Adversarial Generation of Safety-Critical Driving Scenarios

Yuan Gao, Mattia Piccinini, Korbinian Moller, Amr Alanwar, Johannes Betz

机构 * Professorship of Autonomous Vehicle Systems, TUM School of Engineering and Design, Technical University of Munich(自主车辆系统教授职位,技术大学慕尼黑工程与设计学院) Munich Institute of Robotics and Machine Intelligence (MIRMI)(慕尼黑机器人与机器智能研究所) TUM School of Computation, Information and Technology, Department of Computer Engineering, Technical University of Munich(技术大学慕尼黑计算、信息与技术学院,计算机工程系)

专题命中 安全评测 :safety(title,abstract);分类 cs.CL、cs.AI

Comments Final Version and Paper Accepted at IEEE ITSC 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.09816 2025-07-21 cs.CV 78%

Mind the Modality Gap: Towards a Remote Sensing Vision-Language Model via Cross-modal Alignment

Angelos Zavras, Dimitrios Michail, Begüm Demir, Ioannis Papoutsis

机构 * organization= Orion Lab, National Observatory of Athens \& National Technical University of Athens , country= Greece organization= Department of Informatics \& Telematics, Harokopio University of Athens , country= Greece organization= Faculty of Electrical Engineering organization= BIFOLD - Berlin Institute for the Foundations of Learning

专题命中 安全评测 :alignment(title,abstract)

Comments Accepted at the ISPRS Journal of Photogrammetry and Remote Sensing. Our code implementation and weights for all experiments are publicly available at https://github.com/Orion-AI-Lab/MindTheModalityGap

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.13524 2025-07-21 cs.HC cs.AI cs.CY 76%

Humans learn to prefer trustworthy AI over human partners

Yaomin Jiang, Levin Brinkmann, Anne-Marie Nussberger, Ivan Soraperra, Jean-François Bonnefon, Iyad Rahwan

机构 * Toulouse School of Economics, Centre National de la Recherche Scientifique (TSM-R), Université Toulouse Capitole(图卢兹经济学院,法国国家科学研究中心(TSM-R),图卢兹大学)

专题命中 安全评测 :trustworthy(title);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14097 2025-07-21 cs.AI cs.CV 70%

Generative AI-Driven High-Fidelity Human Motion Simulation

Hari Iyer, Neel Macwan, Atharva Jitendra Hude, Heejin Jeong, Shenghan Guo

机构 * The Polytechnic School, Ira A. Fulton Schools of Engineering, Arizona State University(亚利桑那州立大学工程学院Polytechnic学院) School of Manufacturing Systems and Networks, Ira A. Fulton Schools of Engineering, Arizona State University(亚利桑那州立大学工程学院制造系统与网络学院)

专题命中 安全评测 :alignment(abstract);safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.08208 2025-07-21 cs.CL cs.AI 62%

ASTRID -- An Automated and Scalable TRIaD for the Evaluation of RAG-based Clinical Question Answering Systems

Mohita Chowdhury, Yajie Vera He, Jared Joselowitz, Aisling Higham, Ernest Lim

机构 * Ufonia Limited(乌菲尼亚有限公司) University of York(约克大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

Comments 29 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.13390 2025-07-21 cs.CL cs.LG 62%

PARAM-1 BharatGen 2.9B Model

Kundeshwar Pundalik, Piyush Sawarkar, Nihar Sahoo, Abhishek Shinde, Prateek Chanda, Vedant Goswami, Ajay Nagpal, Atul Singh, Viraj Thakur, Vijay Dewane, Aamod Thakur, Bhargav Patel, Smita Gautam, Bhagwan Panditi, Shyam Pawar, Madhav Kotcha, Suraj Racha, Saral Sureka, Pankaj Singh, Rishi Bal, Rohit Saluja, Ganesh Ramakrishnan

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.12674 2025-07-21 cs.CY cs.AI cs.SE 62%

ParaStudent: Generating and Evaluating Realistic Student Code by Teaching LLMs to Struggle

Mihran Miroyan, Rose Niousha, Joseph E. Gonzalez, Gireeja Ranade, Narges Norouzi

机构 * University of California, Berkeley(加州大学伯克利分校)

专题命中 安全评测 :alignment(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.14506 2025-07-21 cs.CV cs.AI cs.CL 62%

On Pre-training of Multimodal Language Models Customized for Chart Understanding

Wan-Cyuan Fan, Yen-Chun Chen, Mengchen Liu, Lu Yuan, Leonid Sigal

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

Comments NeurIPS 2024 Workshop on Adaptive Foundation Models

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14107 2025-07-21 cs.AI cs.IR 57%

Automated Interpretation of Non-Destructive Evaluation Contour Maps Using Large Language Models for Bridge Condition Assessment

Viraj Nishesh Darji, Callie C. Liao, Duoduo Liao

机构 * School of Computing(计算学院) George Mason University(乔治·玛莎大学) College of Science(科学学院)

专题命中 安全评测 :safety(abstract);分类 cs.AI

Journal ref IEEE BigData, Year: 2024; Page: 3258-3263

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.13666 2025-07-21 cs.CL 57%

KiC: Keyword-inspired Cascade for Cost-Efficient Text Generation with LLMs

Woo-Chan Kim, Ji-Hoon Park, Seong-Whan Lee

机构 * Department of Artificial Intelligence, Korea University(人工智能系,韩国大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.17735 2025-07-21 cs.AI 57%

SafeAgent: Safeguarding LLM Agents via an Automated Risk Simulator

Xueyang Zhou, Weidong Wang, Lin Lu, Jiawen Shi, Guiyao Tie, Yongtian Xu, Lixing Chen, Pan Zhou, Neil Zhenqiang Gong, Lichao Sun

专题命中 安全评测 :safety(abstract);分类 cs.AI

Comments 38 pages;12 figures;12 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.13499 2025-07-21 cs.SE cs.AI cs.PL 57%

AI-Assisted Fixes to Code Review Comments at Scale

Chandra Maddila, Negar Ghorbani, James Saindon, Parth Thakkar, Vijayaraghavan Murali, Rui Abreu, Jingyue Shen, Brian Zhou, Nachiappan Nagappan, Peter C. Rigby

机构 * Meta

专题命中 安全评测 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.13981 2025-07-21 cs.CV 50%

Evaluation of Human Visual Privacy Protection: A Three-Dimensional Framework and Benchmark Dataset

Sara Abdulaziz, Giacomo D'Amicantonio, Egor Bondarev

机构 * Eindhoven University of Technology(埃因霍温理工大学)

专题命中 安全评测 :alignment(abstract)

Comments accepted at ICCV'25 workshop CV4BIOM

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.04834 2025-07-21 cs.SE 50%

LLM-Based Multi-Agent Systems for Software Engineering: Literature Review, Vision and the Road Ahead

Junda He, Christoph Treude, David Lo

专题命中 安全评测 :trustworthy(abstract)

Comments TOSEM 2030 Special Issue

详情

展开后加载摘要…

URL PDF HTML 收藏

5. AI治理与伦理 1 篇

2507.13616 2025-07-21 cs.HC cs.CY cs.ET cs.IT cs.MA math.IT 57%

From Firms to Computation: AI Governance and the Evolution of Institutions

Michael S. Harre

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CY

Comments 44 pages

详情

展开后加载摘要…

URL PDF HTML 收藏

6. 其他安全 5 篇

2210.16575 2025-07-21 cs.AI cs.LG cs.RO 81%

Self-Improving Safety Performance of Reinforcement Learning Based Driving with Black-Box Verification Algorithms

Resul Dagdanov, Halil Durmus, Nazim Kemal Ure

机构 * ITU Artificial Intelligence and Data Science Research Center(伊斯坦布尔技术大学人工智能与数据科学研究中心) Department of Aeronautical Engineering(航空工程系) Eatron Technologies(Eatron技术公司) Department of Electronics and Communication Engineering(电子与通信工程系) ITU Artificial Intelligence and Data Science Application and Research Center(伊斯坦布尔技术大学人工智能与数据科学应用与研究中心) Department of Computer Engineering(计算机工程系)

专题命中 其他安全 :safety(title,abstract);分类 cs.AI、cs.LG

Comments 7 pages, 7 figures, 2 tables, published in IEEE International Conference on Robotics and Automation (ICRA), June 2, 2023, London, UK

Journal ref IEEE International Conference on Robotics and Automation (ICRA), 2023, pp. 5631-5637

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.13678 2025-07-21 eess.SY cs.SY 78%

Minimum Clustering of Matrices Based on Phase Alignment

Honghao Wu, Kemi Ding, Li Qiu

专题命中 其他安全 :alignment(title,abstract)

Comments This work has been received by CDC2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14079 2025-07-21 cs.CL cs.AI cs.IR cs.LG 67%

DENSE: Longitudinal Progress Note Generation with Temporal Modeling of Heterogeneous Clinical Notes Across Hospital Visits

Garapati Keerthana, Manik Gupta

机构 * Department of Computer Science and Information Systems, Birla Institute of Technology and Science, Pilani, Hyderabad Campus(计算机科学与信息系统系,比拉理工学院,比拉理工学院,海得拉巴校区)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.12950 2025-07-21 cs.LG 57%

Insights into a radiology-specialised multimodal large language model with sparse autoencoders

Kenza Bouzid, Shruthi Bannur, Felix Meissen, Daniel Coelho de Castro, Anton Schwaighofer, Javier Alvarez-Valle, Stephanie L. Hyland

机构 * Microsoft Research, Health Futures, Cambridge, United Kingdom(微软研究院,健康未来,剑桥,英国)

专题命中 其他安全 :safety(abstract);分类 cs.LG

Comments Actionable Interpretability Workshop at ICML 2025. 24 pages, 7 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.04242 2025-07-21 eess.SY cs.SY 50%

Signal Temporal Logic Control Synthesis among Uncontrollable Dynamic Agents with Conformal Prediction

Xinyi Yu, Yiqi Zhao, Xiang Yin, Lars Lindemann

专题命中 其他安全 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏