arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 8017 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 其他安全 8017 篇

2410.00003 2025-10-21 cs.CV 78%

Large Language Model-Guided Semantic Alignment for Human Activity Recognition

Hua Yan, Heng Tan, Yi Ding, Pengfei Zhou, Vinod Namboodiri, Yu Yang

机构 * Lehigh University(莱维理工大学) University of Texas at Dallas(德克萨斯大学达拉斯分校) University of Pittsburgh(匹兹堡大学)

专题命中 其他安全 :alignment(title);safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12276 2025-10-20 cs.RO 78%

Spatial Forcing: Implicit Spatial Representation Alignment for Vision-language-action Model

Fuhao Li, Wenxuan Song, Han Zhao, Jingbo Wang, Pengxiang Ding, Donglin Wang, Long Zeng, Haoang Li

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州)) Tsinghua University(清华大学) Westlake University(西湖大学) Zhejiang University(浙江大学) South China University of Technology(华南理工大学)

专题命中 其他安全 :alignment(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15693 2025-10-17 cs.CV cs.MM 78%

SCENEFORGE: Enhancing 3D-text alignment with Structured Scene Compositions

Cristian Sbrolli, Matteo Matteucci

机构 * Department of Electronics, Information and Bioengineering(电子、信息与生物工程系)

专题命中 其他安全 :alignment(title,abstract)

Comments to appear in NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.13558 2025-10-16 cs.SD 78%

Steer-MoE: Efficient Audio-Language Alignment with a Mixture-of-Experts Steering Module

Ruitao Feng, Bixi Zhang, Sheng Liang, Zheng Yuan

机构 * The University of Hong Kong, Fauclty of Science, Hong Kong(香港大学科学学院) Aix-Marseille University, Laboratoire Parole et Langage (LPL), France(艾克斯-马赛大学语言与言语实验室(LPL))

专题命中 其他安全 :alignment(title,abstract)

Comments 5 pages, 1 figures. Code is available at: https://github.com/forfrt/SteerMoE. Submitted to ICASSP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.13131 2025-10-16 cs.CV cs.MM 78%

OS-HGAdapter: Open Semantic Hypergraph Adapter for Large Language Models Assisted Entropy-Enhanced Image-Text Alignment

Rongjun Chen, Chengsi Yao, Jinchang Ren, Xianxian Zeng, Peixian Wang, Jun Yuan, Jiawen Li, Huimin Zhao, Xu Lu

机构 * School of Computer Science, Guangdong Polytechnic Normal University(广东 polytechnic 正规大学计算机学院)

专题命中 其他安全 :alignment(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11907 2025-10-15 cs.CV 78%

Task-Specific Dual-Model Framework for Comprehensive Traffic Safety Video Description and Analysis

Blessing Agyei Kyem, Neema Jakisa Owor, Andrews Danyo, Joshua Kofi Asamoah, Eugene Denteh, Tanner Muturi, Anthony Dontoh, Yaw Adu-Gyamfi, Armstrong Aboah

机构 * North Dakota State University(北达科塔州立大学) University of Missouri–Columbia(密苏里大学哥伦比亚分校) University of Memphis(孟菲斯大学)

专题命中 其他安全 :safety(title,abstract)

Comments This paper was accepted at ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.07979 2025-10-13 cs.CV 78%

Visual Representation Alignment for Multimodal Large Language Models

Heeji Yoon, Jaewoo Jung, Junwan Kim, Hyungyu Choi, Heeseong Shin, Sangbeom Lim, Honggyu An, Chaehyun Kim, Jisang Han, Donghyun Kim, Chanho Eom, Sunghwan Hong, Seungryong Kim

机构 * KAIST AI(韩国科学技术院人工智能研究所) New York University(纽约大学) Chung-Ang University(Chung-Ang 大学) Korea University(韩国大学) ETH Zürich(苏黎世联邦理工学院)

专题命中 其他安全 :alignment(title,abstract)

Comments Project Page: https://cvlab-kaist.github.io/VIRAL/

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.06554 2025-10-09 q-bio.QM 78%

UniOTalign: A Global Matching Framework for Protein Alignment via Optimal Transport

Yue Hu, Zanxia Cao, Yingchao Liu

专题命中 其他安全 :alignment(title,abstract)

Comments 7 pages, 1 table

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21665 2025-09-29 cs.HC 78%

Alignment Without Understanding: A Message- and Conversation-Centered Approach to Understanding AI Sycophancy

Lihua Du, Xing Lyu, Lezi Xie, Bo Feng

专题命中 其他安全 :alignment(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.20891 2025-09-26 cs.SD 78%

AIBA: Attention-based Instrument Band Alignment for Text-to-Audio Diffusion

Junyoung Koh, Soo Yong Kim, Gyu Hyeong Choi, Yongwon Choi

机构 * Department of Artificial Intelligence, Yonsei University(人工智能系,延世大学) MAAP LAB, MODULABS(MODULABS 音频实验室) KRAFTON AI Matics Department of Media Software, Sungkyul University(媒体软件系,松谷大学)

专题命中 其他安全 :alignment(title,abstract)

Comments NeurIPS 2025 AI for Music Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.13839 2025-09-18 cs.RO 78%

Pre-Manipulation Alignment Prediction with Parallel Deep State-Space and Transformer Models

Motonari Kambara, Komei Sugiura

机构 * Keio University(keio大学)

专题命中 其他安全 :alignment(title,abstract)

Comments Published in Advanced Robotics

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.20030 2025-09-18 cs.GT cs.DS 78%

Polynomial-Time Approximation Schemes via Utility Alignment: Unit-Demand Pricing and More

Robin Bowers, Marius Garbea, Emmanouil Pountourakis, Samuel Taggart

专题命中 其他安全 :alignment(title,abstract)

Comments To appear in FOCS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.11813 2025-09-08 cs.CV 78%

SEA: Supervised Embedding Alignment for Token-Level Visual-Textual Integration in MLLMs

Yuanyang Yin, Yaqi Zhao, Yajie Zhang, Yuanxing Zhang, Ke Lin, Jiahao Wang, Xin Tao, Pengfei Wan, Wentao Zhang, Feng Zhao

机构 * University of Science and Technology of China(中国科学技术大学) Peking University(北京大学) Kuaishou Technology(快手科技)

专题命中 其他安全 :alignment(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.11171 2025-09-04 eess.SY cs.SY 78%

Preventing Inactive CBF Safety Filters Caused by Invalid Relative Degree Assumptions

Lukas Brunke, Siqi Zhou, Angela P. Schoellig

专题命中 其他安全 :safety(title,abstract)

Comments 8 pages, 4 figures, accepted for publication in the IEEE Transactions on Automatic Control

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00131 2025-09-03 cs.CV eess.IV q-bio.QM 78%

Self-supervised large-scale kidney abnormality detection in drug safety assessment studies

Ivan Slootweg, Natalia P. García-De-La-Puente, Geert Litjens, Salma Dammak

机构 * Department of Pathology, Radboud University Medical Center(拉德堡德大学医学中心病理科部) Instituto Universitario de Investigación en Tecnología Centrada en el Ser Humano, Universitat Politècnica de València(人类中心技术大学研究机构,巴塞罗那理工大学)

专题命中 其他安全 :safety(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20465 2025-08-29 q-bio.NC 78%

On the possibility of deep alignment

Alex B. Kiefer

专题命中 其他安全 :alignment(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.16801 2025-08-27 cs.CV 78%

Decoupled Global-Local Alignment for Improving Compositional Understanding

Xiaoxing Hu, Kaicheng Yang, Jun Wang, Haoran Xu, Ziyong Feng, Yupei Wang

机构 * Beijing Institute of Technology(北京理工大学) Zhejiang University(浙江大学)

专题命中 其他安全 :alignment(title,abstract)

Comments ACMMM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.09963 2025-08-14 eess.SY cs.MA cs.RO cs.SY 78%

Online Safety under Multiple Constraints and Input Bounds using gatekeeper: Theory and Applications

Devansh R. Agrawal, Dimitra Panagou

机构 * Robotics Department, University of Michigan(密歇根大学机器人系)

专题命中 其他安全 :safety(title,abstract)

Comments 6 pages, 2 figures. Accepted for publication in IEEE L-CSS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06300 2025-08-11 cs.HC 78%

Automatic Semantic Alignment of Flow Pattern Representations for Exploration with Large Language Models

Weihan Zhang, Jun Tao

专题命中 其他安全 :alignment(title,abstract)

Comments Accepted by IEEE VIS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01602 2025-08-07 cs.CV 78%

Enhancing Zero-Shot Brain Tumor Subtype Classification via Fine-Grained Patch-Text Alignment

Lubin Gan, Jing Zhang, Linhao Qu, Yijun Wang, Siying Wu, Xiaoyan Sun

机构 * University of Science and Technology of China(科学技术大学) Fudan University(复旦大学)

专题命中 其他安全 :alignment(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01279 2025-08-05 cs.HC 78%

ViseGPT: Towards Better Alignment of LLM-generated Data Wrangling Scripts and User Prompts

Jiajun Zhu, Xinyu Cheng, Zhongsu Luo, Yunfan Zhou, Xinhuan Shu, Di Weng, Yingcai Wu

专题命中 其他安全 :alignment(title,abstract)

Comments Accepted at Annual ACM Symposium on User Interface Software and Technology (UIST'25), September 28-October 1, 2025, Busan, Republic of Korea

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.00945 2025-08-05 cs.CV 78%

Optimizing Vision-Language Consistency via Cross-Layer Regional Attention Alignment

Yifan Wang, Hongfeng Ai, Quangao Liu, Maowei Jiang, Ruiyuan Kang, Ruiqi Li, Jiahua Dong, Mengting Xiao, Cheng Jiang, Chenzhong Li

机构 * School of Medicine, The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)医学院) Shenyang Institute of Automation, Chinese Academy of Sciences(中国科学院沈阳自动化研究所) Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院) Wave and Machine Intelligence Department, Technology Innovation Institute(技术创新研究院波浪与机器智能部门) University of the Chinese Academy of Sciences(中国科学院大学) McGill University(麦吉尔大学)

专题命中 其他安全 :alignment(title,abstract)

Comments 10 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.16772 2025-08-05 cs.CY cs.AI cs.LG 78%

Assessing Social Alignment: Do Personality-Prompted Large Language Models Behave Like Humans?

Ivan Zakazov, Mikolaj Boronski, Lorenzo Drudi, Robert West

机构 * EPFL(苏黎世联邦理工学院)

专题命中 其他安全 :alignment(title);分类 cs.AI、cs.CY、cs.LG

Comments Accepted to NeurIPS 2024 Workshop on Behavioral Machine Learning

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.22511 2025-07-31 cs.ET 78%

Green Wave as an Integral Part for the Optimization of Traffic Efficiency and Safety: A Survey

Kranthi Kumar Talluri, Christopher Stang, Galia Weidl

专题命中 其他安全 :safety(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.13678 2025-07-21 eess.SY cs.SY 78%

Minimum Clustering of Matrices Based on Phase Alignment

Honghao Wu, Kemi Ding, Li Qiu

专题命中 其他安全 :alignment(title,abstract)

Comments This work has been received by CDC2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.11316 2025-07-16 cs.CL cs.AI cs.LG 78%

Internal Value Alignment in Large Language Models through Controlled Value Vector Activation

Haoran Jin, Meng Li, Xiting Wang, Zhihao Xu, Minlie Huang, Yantao Jia, Defu Lian

机构 * University of Science and Technology of China(中国科学技术大学) State Key Laboratory of Cognitive Intelligence(认知智能国家重点实验室) Gaoling School of Artificial Intelligence(光明人工智能学院) Renmin University of China(中国人民大学) Beijing Key Laboratory of Research on Large Models and Intelligent Governance(北京大型模型与智能治理重点实验室) Engineering Research Center of Next-Generation Intelligent Search and Recommendation, MOE(下一代智能搜索与推荐工程技术研究中心) Tsinghua University(清华大学) Huawei Technologies Co. Ltd(华为技术有限公司)

专题命中 其他安全 :alignment(title);分类 cs.CL、cs.AI、cs.LG

Comments 25 pages, 14 figures. Accepted by ACL 2025 (main conference)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.11042 2025-07-16 cs.IR 78%

Aligned Query Expansion: Efficient Query Expansion for Information Retrieval through LLM Alignment

Adam Yang, Gustavo Penha, Enrico Palumbo, Hugues Bouchard

专题命中 其他安全 :alignment(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.11003 2025-07-16 cs.CV 78%

Bridge Feature Matching and Cross-Modal Alignment with Mutual-filtering for Zero-shot Anomaly Detection

Yuhu Bai, Jiangning Zhang, Yunkang Cao, Guangyuan Lu, Qingdong He, Xiangtai Li, Guanzhong Tian

机构 * Zhejiang University(浙江大学) YouTu Lab, Tencent(腾讯YouTu实验室) Huazhong University of Science and Technology(华中科技大学) Peking University(北京大学)

专题命中 其他安全 :alignment(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.09615 2025-07-15 cs.CV 78%

Towards Fine-Grained Adaptation of CLIP via a Self-Trained Alignment Score

Eman Ali, Sathira Silva, Chetan Arora, Muhammad Haris Khan

机构 * Mohamed Bin Zayed University of Artificial Intelligence(莫莫德·本·扎耶德人工智能大学) IIT Delhi(德里印度理工学院) Alexandria University(亚历山大大学)

专题命中 其他安全 :alignment(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.07568 2025-07-11 stat.ME eess.IV 78%

Learnable Retrieval Enhanced Visual-Text Alignment and Fusion for Radiology Report Generation

Qin Zhou, Guoyan Liang, Xindi Li, Jingyuan Chen, Wang Zhe, Chang Yao, Sai Wu

专题命中 其他安全 :alignment(title,abstract)

Comments 10 pages,3 figures, conference

详情

展开后加载摘要…

URL PDF HTML 收藏