arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 9419 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 9419 篇

2505.11077 2025-05-30 eess.SY cs.SY 78%

LLM-Enhanced Symbolic Control for Safety-Critical Applications

Amir Bayat, Alessandro Abate, Necmiye Ozay, Raphael M. Jungers

专题命中 安全评测 :safety(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.20136 2025-05-27 cs.SE cs.CR 78%

Engineering Trustworthy Machine-Learning Operations with Zero-Knowledge Proofs

Filippo Scaramuzza, Giovanni Quattrocchi, Damian A. Tamburri

专题命中 安全评测 :trustworthy(title,abstract)

Comments This work has been submitted to the IEEE for possible publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18730 2025-05-27 cs.CV 78%

Align Beyond Prompts: Evaluating World Knowledge Alignment in Text-to-Image Generation

Wenchao Zhang, Jiahe Tian, Runze He, Jizhong Han, Jiao Dai, Miaomiao Feng, Wei Mi, Xiaodan Zhang

机构 * Institute of Information Engineering, Chinese Academy of Sciences(信息工程研究所,中国科学院)

专题命中 安全评测 :alignment(title,abstract)

Comments Code: https://github.com/smile365317/ABP

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.17727 2025-05-26 cs.CV 78%

SafeMVDrive: Multi-view Safety-Critical Driving Video Synthesis in the Real World Domain

Jiawei Zhou, Linye Lyu, Zhuotao Tian, Cheng Zhuo, Yu Li

专题命中 安全评测 :safety(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.10917 2025-05-20 cs.CV 78%

VISTA: Enhancing Vision-Text Alignment in MLLMs via Cross-Modal Mutual Information Maximization

Mingxiao Li, Na Su, Fang Qu, Zhizhou Zhong, Ziyang Chen, Yuan Li, Zhaopeng Tu, Xiaolong Li

机构 * Tencent(腾讯)

专题命中 安全评测 :alignment(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.11845 2025-05-20 cs.CV 78%

ElderFallGuard: Real-Time IoT and Computer Vision-Based Fall Detection System for Elderly Safety

Tasrifur Riahi, Md. Azizul Hakim Bappy, Md. Mehedi Islam

机构 * Institute of Information and Communicaton Technology, Bangladesh University of Engineering Technology(信息与通信技术学院,孟加拉国工程科技大学) Dept. of Electronics and Communication Engineering, Hajee Mohammad Danesh Science and Technology University(电子与通信工程系,海杰穆罕默德丹尼什科学与技术大学)

专题命中 安全评测 :safety(title,abstract)

Comments 9 page, 1 table, 5 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.13707 2025-05-13 eess.SY cs.RO cs.SY 78%

Safety-Critical Formation Control of Non-Holonomic Multi-Robot Systems in Communication-Limited Environments

Vishrut Bohara, Siavash Farzan

机构 * Robotics Engineering Department, Worcester Polytechnic Institute(沃斯彻斯特理工学院机器人工程系) Electrical Engineering Department, California Polytechnic State University(加州州立大学弗雷斯诺分校电子工程系)

专题命中 安全评测 :safety(title,abstract)

Comments Under review. Video demonstration: https://vimeo.com/1075016147/41612f2f8c

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.02598 2025-05-06 cs.RO cs.SY eess.SY 78%

LiDAR-Inertial SLAM-Based Navigation and Safety-Oriented AI-Driven Control System for Skid-Steer Robots

Mehdi Heydari Shahna, Eemil Haaparanta, Pauli Mustalahti, Jouni Mattila

机构 * Faculty of Engineering and Natural Sciences, Tampere University(工程与自然科学学院,塔尔库大学)

专题命中 安全评测 :safety(title,abstract)

Comments This paper has been submitted in the IEEE CDC 2025 for potential presentation

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.17968 2025-04-28 cs.RO 78%

Virtual Roads, Smarter Safety: A Digital Twin Framework for Mixed Autonomous Traffic Safety Analysis

Hao Zhang, Ximin Yue, Kexin Tian, Sixu Li, Keshu Wu, Zihao Li, Dominique Lord, Yang Zhou

机构 * Zachry Department of Civil and Environmental Engineering, Texas A&M University(土木与环境工程系,德克萨斯A&M大学) Department of Landscape Architecture and Urban Planning(景观建筑与城市规划系)

专题命中 安全评测 :safety(title,abstract)

Comments 14 pages, 18 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.12018 2025-04-17 cs.CV 78%

Instruction-augmented Multimodal Alignment for Image-Text and Element Matching

Xinli Yue, JianHui Sun, Junda Lu, Liangchao Yao, Fan Xia, Tianyi Wang, Fengyun Rao, Jing Lyu, Yuetang Deng

专题命中 安全评测 :alignment(title,abstract)

Comments Accepted to CVPR 2025 Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.09757 2025-04-15 cs.CR 78%

Alleviating the Fear of Losing Alignment in LLM Fine-tuning

Kang Yang, Guanhong Tao, Xun Chen, Jun Xu

专题命中 安全评测 :alignment(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.07556 2025-04-11 cs.CV 78%

TokenFocus-VQA: Enhancing Text-to-Image Alignment with Position-Aware Focus and Multi-Perspective Aggregations on LVLMs

Zijian Zhang, Xuhui Zheng, Xuecheng Wu, Chong Peng, Xuezhi Cao

专题命中 安全评测 :alignment(title,abstract)

Comments 10 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.02406 2025-04-04 cs.NI 78%

Lifecycle Management of Trustworthy AI Models in 6G Networks: The REASON Approach

Juan Parra-Ullauri, Xueqing Zhou, Shadi Moazzeni, Rasheed Hussain, Xenofon Vasilakos, Yulei Wu, Renjith Baby, M M Hassan Mahmud, Gabriele Incorvaia, Darryl Hond, Hamid Asgari, Andrea Tassi, Daniel Warren, Dimitra Simeonidou

专题命中 安全评测 :trustworthy(title,abstract)

Journal ref IEEE Wireless Communications, vol. 32, no. 2, pp. 42-51, April 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.02017 2025-04-04 cs.SE 78%

Enhancing LLMs in Long Code Translation through Instrumentation and Program State Alignment

Li Xin-Ye, Du Ya-Li, Li Ming

专题命中 安全评测 :alignment(title,abstract)

Comments 20 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.18609 2025-03-28 cs.CV 78%

Video-Panda: Parameter-efficient Alignment for Encoder-free Video-Language Models

Jinhui Yi, Syed Talal Wasim, Yanan Luo, Muzammal Naseer, Juergen Gall

专题命中 安全评测 :alignment(title,abstract)

Comments CVPR 2025 camera-ready version

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.03534 2025-03-26 cs.SE cs.SY eess.SY 78%

Simulation-Based Application of Safety of The Intended Functionality to Mitigate Foreseeable Misuse in Automated Driving Systems

Milin Patel, Rolf Jung

专题命中 安全评测 :safety(title,abstract)

Comments SAE MobilityRxiv Preprint, 2023

Journal ref SAE MobilityRxiv Preprint, 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.15522 2025-03-21 cs.HC 78%

"I don't like things where I do not have control": Participants' Experience of Trustworthy Interaction with Autonomous Vehicles

Ana Tanevska, Katie Winkle, Ginevra Castellano

专题命中 安全评测 :trustworthy(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.09594 2025-03-13 cs.CV cs.RO 78%

SimLingo: Vision-Only Closed-Loop Autonomous Driving with Language-Action Alignment

Katrin Renz, Long Chen, Elahe Arani, Oleg Sinavski

专题命中 安全评测 :alignment(title,abstract)

Comments CVPR 2025. 1st Place @ CARLA Challenge 2024. Challenge tech report (preliminary version of SimLingo): arXiv:2406.10165

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.11087 2025-03-05 cs.CV 78%

Locality Alignment Improves Vision-Language Models

Ian Covert, Tony Sun, James Zou, Tatsunori Hashimoto

专题命中 安全评测 :alignment(title,abstract)

Comments ICLR 2025 Camera-Ready

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.18883 2025-03-04 cs.SE 78%

Towards More Trustworthy Deep Code Models by Enabling Out-of-Distribution Detection

Yanfu Yan, Viet Duong, Huajie Shao, Denys Poshyvanyk

专题命中 安全评测 :trustworthy(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.18757 2025-02-27 cs.IR 78%

Training Large Recommendation Models via Graph-Language Token Alignment

Mingdai Yang, Zhiwei Liu, Liangwei Yang, Xiaolong Liu, Chen Wang, Hao Peng, Philip S. Yu

专题命中 安全评测 :alignment(title,abstract)

Comments 5 pages. Accepted by www'25 as short paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.00809 2025-02-26 cs.LG cs.AI cs.CL 78%

Adaptive Segment-level Reward: Bridging the Gap Between Action and Reward Space in Alignment

Yanshi Li, Shaopan Xiong, Gengru Chen, Xiaoyang Li, Yijia Luo, Xingyuan Bu, Yingshui Tan, Wenbo Su, Bo Zheng

专题命中 安全评测 :alignment(title);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.14881 2025-02-24 cs.CR cs.CV 78%

A Survey of Safety on Large Vision-Language Models: Attacks, Defenses and Evaluations

Mang Ye, Xuankun Rong, Wenke Huang, Bo Du, Nenghai Yu, Dacheng Tao

专题命中 安全评测 :safety(title,abstract)

Comments 22 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.02097 2025-02-23 cs.CR 78%

DomainHarvester: Harvesting Infrequently Visited Yet Trustworthy Domain Names

Daiki Chiba, Hiroki Nakano, Takashi Koide

专题命中 安全评测 :trustworthy(title,abstract)

Comments Originally presented at IEEE CCNC 2025. An extended version of this work has been published in IEEE Access: https://doi.org/10.1109/ACCESS.2025.3539882

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.04076 2025-02-07 cs.CV 78%

Content-Rich AIGC Video Quality Assessment via Intricate Text Alignment and Motion-Aware Consistency

Shangkun Sun, Xiaoyu Liang, Bowen Qu, Wei Gao

专题命中 安全评测 :alignment(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.15363 2025-01-28 cs.CR cs.CV 78%

AI-Driven Secure Data Sharing: A Trustworthy and Privacy-Preserving Approach

Al Amin, Kamrul Hasan, Sharif Ullah, Liang Hong

专题命中 安全评测 :trustworthy(title,abstract)

Comments 6 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.09102 2025-01-17 cs.SI cs.AI cs.CY cs.LG 78%

Tracking the Takes and Trajectories of English-Language News Narratives across Trustworthy and Worrisome Websites

Hans W. A. Hanley, Emily Okabe, Zakir Durumeric

专题命中 安全评测 :trustworthy(title);分类 cs.AI、cs.CY、cs.LG

Comments To appear at USENIX Security Symposium 2025. Keywords: Misinformation, News, Narratives, LLMs, Stance-Detection

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.08231 2024-12-24 cs.IR 78%

DaRec: A Disentangled Alignment Framework for Large Language Model and Recommender System

Xihong Yang, Heming Jing, Zixing Zhang, Jindong Wang, Huakang Niu, Shuaiqiang Wang, Yu Lu, Junfeng Wang, Dawei Yin, Xinwang Liu, En Zhu, Defu Lian, Erxue Min

专题命中 安全评测 :alignment(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.06296 2024-12-10 cs.SD eess.AS 78%

VidMusician: Video-to-Music Generation with Semantic-Rhythmic Alignment via Hierarchical Visual Features

Sifei Li, Binxin Yang, Chunji Yin, Chong Sun, Yuxin Zhang, Weiming Dong, Chen Li

专题命中 安全评测 :alignment(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.04363 2024-12-06 cs.HC 78%

Challenges in Trustworthy Human Evaluation of Chatbots

Wenting Zhao, Alexander M. Rush, Tanya Goyal

专题命中 安全评测 :trustworthy(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏