arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 9419 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 9419 篇

2510.02987 2025-10-06 cs.CV 78%

TIT-Score: Evaluating Long-Prompt Based Text-to-Image Alignment via Text-to-Image-to-Text Consistency

Juntong Wang, Huiyu Duan, Jiarui Wang, Ziheng Jia, Guangtao Zhai, Xiongkuo Min

机构 * Institute of Image Communication and Network Engineering(图像通信与网络工程研究所) MoE Key Lab of Artificial Intelligence, AI Institute(人工智能关键实验室) Shanghai Jiao Tong University(上海交通大学)

专题命中 安全评测 :alignment(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.26122 2025-10-01 math.NA cs.NA 78%

Trustworthy AI in numerics: On verification algorithms for neural network-based PDE solvers

Emil Haugen, Alexei Stepanenko, Anders C. Hansen

专题命中 安全评测 :trustworthy(title,abstract)

Comments 25 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.04974 2025-10-01 eess.AS cs.SD 78%

From Voice to Safety: Language AI Powered Pilot-ATC Communication Understanding for Airport Surface Movement Collision Risk Assessment

Yutian Pang, Andrew Paul Kendall, Alex Porcayo, Mariah Barsotti, Anahita Jain, John-Paul Clarke

机构 * Department of Aerospace Engineering and Engineering Mechanics(航空航天工程与工程力学系)

专题命中 安全评测 :safety(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23760 2025-09-30 cs.CV 78%

UniAlignment: Semantic Alignment for Unified Image Generation, Understanding, Manipulation and Perception

Xinyang Song, Libin Wang, Weining Wang, Shaozhen Liu, Dandan Zheng, Jingdong Chen, Qi Li, Zhenan Sun

机构 * Ant Group(蚂蚁集团)

专题命中 安全评测 :alignment(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.20393 2025-09-26 cs.CY cs.AI cs.LG 78%

The Secret Agenda: LLMs Strategically Lie and Our Current Safety Tools Are Blind

Caleb DeLeeuw, Gaurav Chawla, Aniket Sharma, Vanessa Dietze

机构 * Independent Researcher(独立研究者)

专题命中 安全评测 :safety(title);分类 cs.AI、cs.CY、cs.LG

Comments 9 pages plus citations and appendix, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.14446 2025-09-19 q-bio.NC 78%

Mouse vs. AI: A Neuroethological Benchmark for Visual Robustness and Neural Alignment

Marius Schneider, Joe Canzano, Jing Peng, Yuchen Hou, Spencer LaVere Smith, Michael Beyeler

专题命中 安全评测 :alignment(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.08997 2025-09-12 cs.HC 78%

YouthSafe: A Youth-Centric Safety Benchmark and Safeguard Model for Large Language Models

Yaman Yu, Yiren Liu, Jacky Zhang, Yun Huang, Yang Wang

专题命中 安全评测 :safety(title,abstract)

Comments 15 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00700 2025-09-10 cs.CV 78%

Prompt the Unseen: Evaluating Visual-Language Alignment Beyond Supervision

Raehyuk Jung, Seungjun Yu, Hyunjung Shim

专题命中 安全评测 :alignment(title,abstract)

Comments Link to publicly available codes is added

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.10610 2025-09-03 cs.RO cs.SY eess.SY 78%

Safety-Critical Human-Machine Shared Driving for Vehicle Collision Avoidance based on Hamilton-Jacobi reachability

Shiyue Zhao, Junzhi Zhang, Rui Zhou, Neda Masoud, Jianxiong Li, Helai Huang, Shijie Zhao

专题命中 安全评测 :safety(title,abstract)

Comments 36 pages, 15 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.14527 2025-08-26 cs.CV 78%

Adversarial Generation and Collaborative Evolution of Safety-Critical Scenarios for Autonomous Vehicles

Jiangfan Liu, Yongkang Guo, Fangzhi Zhong, Tianyuan Zhang, Zonglei Jing, Siyuan Liang, Jiakai Wang, Mingchuan Zhang, Aishan Liu, Xianglong Liu

机构 * Beihang University(北京航空航天大学) Nanyang Technological University(南洋理工大学) Zhongguancun Laboratory(中关村实验室) Henan University of Science and Technology(河南科技大学)

专题命中 安全评测 :safety(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.16213 2025-08-25 cs.CV 78%

MedOmni-45°: A Safety-Performance Benchmark for Reasoning-Oriented LLMs in Medicine

Kaiyuan Ji, Yijin Guo, Zicheng Zhang, Xiangyang Zhu, Yuan Tian, Ning Liu, Guangtao Zhai

专题命中 安全评测 :safety(title,abstract)

Comments 9 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.01408 2025-08-19 cs.RO cs.CV 78%

From Shadows to Safety: Occlusion Tracking and Risk Mitigation for Urban Autonomous Driving

Korbinian Moller, Luis Schwarzmeier, Johannes Betz

专题命中 安全评测 :safety(title,abstract)

Comments 8 Pages. Submitted to the IEEE Intelligent Vehicles Symposium (IV 2025), Romania

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.16867 2025-08-19 cs.CV 78%

ETVA: Evaluation of Text-to-Video Alignment via Fine-grained Question Generation and Answering

Kaisi Guan, Zhengfeng Lai, Yuchong Sun, Peng Zhang, Wei Liu, Kieran Liu, Meng Cao, Ruihua Song

机构 * Renmin University of China(中国人民大学) Apple(苹果公司)

专题命中 安全评测 :alignment(title,abstract)

Comments International Conference on Computer Vision 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.00399 2025-08-15 cs.CV 78%

iSafetyBench: A video-language benchmark for safety in industrial environment

Raiyaan Abdullah, Yogesh Singh Rawat, Shruti Vyas

机构 * University of Central Florida(中央佛罗里达大学)

专题命中 安全评测 :safety(title,abstract)

Comments Accepted to VISION'25 - ICCV 2025 workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07560 2025-08-12 cs.RO cs.CV 78%

Progressive Bird's Eye View Perception for Safety-Critical Autonomous Driving: A Comprehensive Survey

Yan Gong, Naibang Wang, Jianli Lu, Xinyu Zhang, Yongsheng Gao, Jie Zhao, Zifan Huang, Haozhi Bai, Nanxin Zeng, Nayu Su, Lei Yang, Ziying Song, Xiaoxi Hu, Xinmin Jiang, Xiaojuan Zhang, Susanto Rahardja

机构 * State Key Laboratory of Robotics and System(机器人系统国家重点实验室) Harbin Institute of Technology(哈尔滨工业大学) State Key Laboratory of Intelligent Green Vehicle and Mobility(智能绿色车辆与移动性国家重点实验室) Tsinghua University(清华大学) the School of Mechanical and Aerospace Engineering(机械与航空航天工程学院) Nanyang Technological University(南洋理工大学) Beijing Key Laboratory of Traffic Data Mining and Embodied Intelligence(北京交通数据挖掘与具身智能重点实验室) Beijing Jiaotong University(北京交通大学) the Institute for Infocomm Research(信息通信研究所) A*STAR the Engineering Cluster(工程集群) the Singapore Institute of Technology(新加坡理工学院)

专题命中 安全评测 :safety(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.03834 2025-08-11 cs.RO cs.CV 78%

CARE: Enhancing Safety of Visual Navigation through Collision Avoidance via Repulsive Estimation

Joonkyung Kim, Joonyeol Sim, Woojun Kim, Katia Sycara, Changjoo Nam

机构 * Department of Electronic Engineering, Sogang University(电子工程系,首尔大学) Robotics Institute, Carnegie Mellon University(机器人研究所,卡内基梅隆大学)

专题命中 安全评测 :safety(title,abstract)

Comments 16 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04240 2025-08-07 eess.SP 78%

ChineseEEG-2: An EEG Dataset for Multimodal Semantic Alignment and Neural Decoding during Reading and Listening

Sitong Chen, Beiqianyi Li, Cuilin He, Dongyang Li, Mingyang Wu, Xinke Shen, Song Wang, Xuetao Wei, Xindi Wang, Haiyan Wu, Quanying Liu

专题命中 安全评测 :alignment(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.10817 2025-08-01 stat.AP 78%

Is Your Model Risk ALARP? Evaluating Prospective Safety-Critical Applications of Complex Models

Domenic Di Francesco, Alan Forrest, Fiona McGarry, Nicholas Hall, Adam Sobey

专题命中 安全评测 :safety(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.22389 2025-07-31 cs.RO cs.SY eess.SY 78%

Safety Evaluation of Motion Plans Using Trajectory Predictors as Forward Reachable Set Estimators

Kaustav Chakraborty, Zeyuan Feng, Sushant Veer, Apoorva Sharma, Wenhao Ding, Sever Topan, Boris Ivanovic, Marco Pavone, Somil Bansal

机构 * Department of Electrical Engineering, University of Southern California(电气工程系,美国南加州大学) Department of Aeronautics and Astronautics, Stanford University(航空与宇航系,斯坦福大学) NVIDIA Research(NVIDIA研究)

专题命中 安全评测 :safety(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21741 2025-07-30 cs.CV cs.MM 78%

MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces

Shaojun E, Yuchen Yang, Jiaheng Wu, Yan Zhang, Tiejun Zhao, Ziyan Chen

机构 * Global Tone Communication Technology Co., Ltd.(全球 tone 通信技术有限公司) Faculty of computing, Harbin Institute of Technology(哈尔滨工业大学计算机学院) School of Computer Science and Technology, Beijing Jiaotong University(北京交通大学计算机科学与技术学院)

专题命中 安全评测 :alignment(title,abstract)

Comments 9 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21145 2025-07-30 cs.CR 78%

Leveraging Trustworthy AI for Automotive Security in Multi-Domain Operations: Towards a Responsive Human-AI Multi-Domain Task Force for Cyber Social Security

Vita Santa Barletta, Danilo Caivano, Gabriel Cellammare, Samuele del Vescovo, Annita Larissa Sciacovelli

专题命中 安全评测 :trustworthy(title,abstract)

Comments 13 pages, 6 figures, 1 table

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.09816 2025-07-21 cs.CV 78%

Mind the Modality Gap: Towards a Remote Sensing Vision-Language Model via Cross-modal Alignment

Angelos Zavras, Dimitrios Michail, Begüm Demir, Ioannis Papoutsis

机构 * organization= Orion Lab, National Observatory of Athens \& National Technical University of Athens , country= Greece organization= Department of Informatics \& Telematics, Harokopio University of Athens , country= Greece organization= Faculty of Electrical Engineering organization= BIFOLD - Berlin Institute for the Foundations of Learning

专题命中 安全评测 :alignment(title,abstract)

Comments Accepted at the ISPRS Journal of Photogrammetry and Remote Sensing. Our code implementation and weights for all experiments are publicly available at https://github.com/Orion-AI-Lab/MindTheModalityGap

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.12377 2025-07-15 cs.CV 78%

Alignment and Adversarial Robustness: Are More Human-Like Models More Secure?

Blaine Hoak, Kunyang Li, Patrick McDaniel

机构 * University of Wisconsin-Madison(威斯康星大学麦迪逊分校)

专题命中 安全评测 :alignment(title,abstract)

Comments Accepted to International Workshop on Security and Privacy-Preserving AI/ML (SPAIML) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.02137 2025-07-11 cs.SE 78%

Towards Trustworthy Sentiment Analysis in Software Engineering: Dataset Characteristics and Tool Selection

Martin Obaidi, Marc Herrmann, Jil Klünder, Kurt Schneider

专题命中 安全评测 :trustworthy(title,abstract)

Comments This paper has been accepted at the RETRAI workshop of the 33rd IEEE International Requirements Engineering Workshop (REW 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.19052 2025-06-25 cs.CR 78%

Trustworthy Artificial Intelligence for Cyber Threat Analysis

Shuangbao Paul Wang, Paul Mullin

专题命中 安全评测 :trustworthy(title,abstract)

Journal ref Springer Lecture Note in Networks and Systems. 978-3-031-16071-4,Vol I, LNNS 542. pp 493-504. 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.18938 2025-06-25 cs.CV cs.SY eess.SY 78%

Bird's-eye view safety monitoring for the construction top under the tower crane

Yanke Wang, Yu Hin Ng, Haobo Liang, Ching-Wei Chang, Hao Chen

机构 * Hong Kong Center for Construction Robotics(香港建设机器人中心) The Hong Kong University of Science and Technology(香港理工大学)

专题命中 安全评测 :safety(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.17812 2025-06-24 cs.SE 78%

Is Your Automated Software Engineer Trustworthy?

Noble Saji Mathews, Meiyappan Nagappan

专题命中 安全评测 :trustworthy(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.13929 2025-06-11 cs.HC 78%

Awake at the Wheel: Enhancing Automotive Safety through EEG-Based Fatigue Detection

Gourav Siddhad, Sayantan Dey, Partha Pratim Roy, Masakazu Iwamura

专题命中 安全评测 :safety(title,abstract)

Comments 7 Pages, 2 Figure, 1 Table

Journal ref International Conference on Pattern Recognition (ICPR) 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.07633 2025-06-10 cs.RO 78%

Blending Participatory Design and Artificial Awareness for Trustworthy Autonomous Vehicles

Ana Tanevska, Ananthapathmanabhan Ratheesh Kumar, Arabinda Ghosh, Ernesto Casablanca, Ginevra Castellano, Sadegh Soudjani

机构 * Uppsala University(乌普萨拉大学) Max Planck Institute for Software Systems(马克斯·普朗克软件系统研究所) Newcastle University(新castle大学)

专题命中 安全评测 :trustworthy(title,abstract)

Comments Submitted to IEEE RO-MAN 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.01667 2025-06-10 cs.CV 78%

VProChart: Answering Chart Question through Visual Perception Alignment Agent and Programmatic Solution Reasoning

Muye Huang, Lingling Zhang, Lai Han, Wenjun Wu, Xinyu Zhang, Jun Liu

专题命中 安全评测 :alignment(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏