arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 8057 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 其他安全 8057 篇

2408.11963 2025-03-10 cs.CV cs.AI 57%

Real-Time Incremental Explanations for Object Detectors in Autonomous Driving

Santiago Calderón-Peña, Hana Chockler, David A. Kelly

机构 * King’s College London(伦敦国王学院)

专题命中 其他安全 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.04252 2025-03-07 cs.DB cs.LG 57%

RCRank: Multimodal Ranking of Root Causes of Slow Queries in Cloud Database Systems

Biao Ouyang, Yingying Zhang, Hanyin Cheng, Yang Shu, Chenjuan Guo, Bin Yang, Qingsong Wen, Lunting Fan, Christian S. Jensen

机构 * East China Normal University(华东师范大学) Alibaba Cloud Computing(阿里巴巴云计算) Squirrel Ai Learning(松鼠AI教育科技) Aalborg University(奥尔堡大学)

专题命中 其他安全 :alignment(abstract);分类 cs.LG

Comments Accepted by VLDB 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.04229 2025-03-07 cs.CV cs.LG 57%

Synthetic Data is an Elegant GIFT for Continual Vision-Language Models

Bin Wu, Wuxuan Shi, Jinqiao Wang, Mang Ye

机构 * School of Computer Science, Wuhan University(武汉大学计算机学院) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) Wuhan AI Research(武汉人工智能研究院) Taikang Center for Life and Medical Sciences, Wuhan University(武汉大学泰康生命医学中心)

专题命中 其他安全 :alignment(abstract);分类 cs.LG

Comments This work is accepted by CVPR 2025. Modifications may be performed

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.03474 2025-03-06 cs.CL 57%

Enhancing Spoken Discourse Modeling in Language Models Using Gestural Cues

Varsha Suresh, M. Hamza Mughal, Christian Theobalt, Vera Demberg

机构 * Saarland University(萨尔大学) Max Planck Institute for Informatics(马克斯·普朗克信息学研究所)

专题命中 其他安全 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.03278 2025-03-06 cs.CV cs.CL 57%

Enhancing Abnormality Grounding for Vision Language Models with Knowledge Descriptions

Jun Li, Che Liu, Wenjia Bai, Rossella Arcucci, Cosmin I. Bercea, Julia A. Schnabel

机构 * Technical University of Munich(慕尼黑工业大学) Munich Center for Machine Learning(慕尼黑机器学习中心) Helmholtz AI(亥姆霍兹人工智能研究所) Helmholtz Munich(亥姆霍兹慕尼黑中心) Imperial College London(伦敦帝国学院) King’s College London(伦敦国王学院)

专题命中 其他安全 :alignment(abstract);分类 cs.CL

Comments 11 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.01275 2025-03-06 cs.CL 57%

Enhancing Non-English Capabilities of English-Centric Large Language Models through Deep Supervision Fine-Tuning

Wenshuai Huo, Xiaocheng Feng, Yichong Huang, Chengpeng Fu, Baohang Li, Yangfan Ye, Zhirui Zhang, Dandan Tu, Duyu Tang, Yunfei Lu, Hui Wang, Bing Qin

专题命中 其他安全 :alignment(abstract);分类 cs.CL

Comments Accepted at AAAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.02967 2025-03-06 cs.CV cs.LG 57%

Revolutionizing Traffic Management with AI-Powered Machine Vision: A Step Toward Smart Cities

Seyed Hossein Hosseini DolatAbadi, Sayyed Mohammad Hossein Hashemi, Mohammad Hosseini, Moein-Aldin AliHosseini

专题命中 其他安全 :safety(abstract);分类 cs.LG

Comments 6 pages, 1 figure, 2 tables, accepted to 1th AITC conference in University Of Isfahan

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.02335 2025-03-05 cs.SE cs.CL 57%

Unlocking a New Rust Programming Experience: Fast and Slow Thinking with LLMs to Conquer Undefined Behaviors

Renshuang Jiang, Pan Dong, Zhenling Duan, Yu Shi, Xiaoxiang Fang, Yan Ding, Jun Ma, Shuai Zhao, Zhe Jiang

机构 * National University of Defense Technology(国防科技大学) Sun Yat-sen University(中山大学) Southeast University(东南大学)

专题命中 其他安全 :safety(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.02239 2025-03-05 cs.AI 57%

V2X-LLM: Enhancing V2X Integration and Understanding in Connected Vehicle Corridors

Keshu Wu, Pei Li, Yang Zhou, Rui Gan, Junwei You, Yang Cheng, Jingwen Zhu, Steven T. Parker, Bin Ran, David A. Noyce, Zhengzhong Tu

机构 * University of Wisconsin-Madison(威斯康星大学麦迪逊分校) Texas A&M University(德克萨斯农工大学)

专题命中 其他安全 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.10038 2025-03-05 cs.AI 57%

POI-Enhancer: An LLM-based Semantic Enhancement Framework for POI Representation Learning

Jiawei Cheng, Jingyuan Wang, Yichuan Zhang, Jiahao Ji, Yuanshao Zhu, Zhibo Zhang, Xiangyu Zhao

专题命中 其他安全 :alignment(abstract);分类 cs.AI

Comments AAAI 25

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.01632 2025-03-04 cs.AI 57%

CoT-VLM4Tar: Chain-of-Thought Guided Vision-Language Models for Traffic Anomaly Resolution

Tianchi Ren, Haibo Hu, Jiacheng Zuo, Xinhong Chen, Jianping Wang, Chun Jason Xue, Jen-Ming Wu, Nan Guan

机构 * City University of Hong Kong(香港城市大学) Soochow University(苏州大学) Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学) Honhai Research Institute(鸿海研究院)

专题命中 其他安全 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.00669 2025-03-04 cs.LG 57%

The Role, Trends, and Applications of Machine Learning in Undersea Communication: A Bangladesh Perspective

Yousuf Islam, Sumon Chandra Das, Md. Jalal Uddin Chowdhury

专题命中 其他安全 :safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.00427 2025-03-04 cs.SD cs.AI eess.AS 57%

Language Model Mapping in Multimodal Music Learning: A Grand Challenge Proposal

Daniel Chin, Gus Xia

机构 * NYU Shanghai(上海纽约大学) MBZUAI(穆罕默德·本·扎耶德人工智能大学)

专题命中 其他安全 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.18821 2025-03-04 cs.LG 57%

CAMEx: Curvature-aware Merging of Experts

Dung V. Nguyen, Minh H. Nguyen, Luc Q. Nguyen, Rachel S. Y. Teo, Tan M. Nguyen, Linh Duy Tran

机构 * Faculty of Mathematics and Informatics, Hanoi University of Science and Technology(河内科学技术大学数学与信息学院) Viettel AI, Viettel Group(Viettel集团Viettel人工智能公司) Department of Mathematics, National University of Singapore(新加坡国立大学数学系)

专题命中 其他安全 :alignment(abstract);分类 cs.LG

Comments 10 pages, 5 Figures, 7 Tables. Published at ICLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.04236 2025-03-04 cs.CV cs.CL 57%

CogCoM: A Visual Language Model with Chain-of-Manipulations Reasoning

Ji Qi, Ming Ding, Weihan Wang, Yushi Bai, Qingsong Lv, Wenyi Hong, Bin Xu, Lei Hou, Juanzi Li, Yuxiao Dong, Jie Tang

机构 * Tsinghua University(清华大学) Zhipu AI(智谱AI)

专题命中 其他安全 :alignment(abstract);分类 cs.CL

Comments 21 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.02447 2025-03-03 cs.LG 57%

PLeaS -- Merging Models with Permutations and Least Squares

Anshul Nasery, Jonathan Hayase, Pang Wei Koh, Sewoong Oh

机构 * University of Washington(华盛顿大学) Allen Institute for AI(艾伦人工智能研究所)

专题命中 其他安全 :alignment(abstract);分类 cs.LG

Comments Accepted to CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.16235 2025-02-28 cs.AI 57%

Dynamic Parallel Tree Search for Efficient LLM Reasoning

Yifu Ding, Wentao Jiang, Shunyu Liu, Yongcheng Jing, Jinyang Guo, Yingjie Wang, Jing Zhang, Zengmao Wang, Ziwei Liu, Bo Du, Xianglong Liu, Dacheng Tao

机构 * Beihang University(北京航空航天大学) Nanyang Technological University(南洋理工大学) Wuhan University(武汉大学)

专题命中 其他安全 :alignment(abstract);分类 cs.AI

Comments 17 pages, 11 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.03006 2025-02-28 cs.LG cond-mat.dis-nn 57%

Formation of Representations in Neural Networks

Liu Ziyin, Isaac Chuang, Tomer Galanti, Tomaso Poggio

机构 * Massachusetts Institute of Technology(麻省理工学院) Texas A&M University(德克萨斯农工大学) NTT Research(NTT研究所)

专题命中 其他安全 :alignment(abstract);分类 cs.LG

Comments ICLR 2025 Spotlight

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.16870 2025-02-28 cs.HC cs.LG 57%

Quantifying Visual Properties of GAM Shape Plots: Impact on Perceived Cognitive Load and Interpretability

Sven Kruschel, Lasse Bohlen, Julian Rosenberger, Patrick Zschech, Mathias Kraus

机构 * University of Regensburg(雷根斯堡大学) Leipzig University(莱比锡大学)

专题命中 其他安全 :alignment(abstract);分类 cs.LG

Comments to be published in proceedings of the 58th Hawaii International Conference on System Sciences (HICSS)

Journal ref Proceedings of the 58th Hawaii International Conference on System Sciences 2025 (HICSS)

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.16233 2025-02-28 cs.LG math.GR 57%

Algebraic Adversarial Attacks on Integrated Gradients

Lachlan Simpson, Federico Costanza, Kyle Millar, Adriel Cheng, Cheng-Chew Lim, Hong Gunn Chew

机构 * School of Electrical and Mechanical Engineering, The University of Adelaide(阿德莱德大学机电工程学院) School of Computer and Mathematical Sciences, The University of Adelaide(阿德莱德大学计算机与数学科学学院) Information Sciences Division, Defence Science & Technology Group(国防科技集团信息科学部)

专题命中 其他安全 :safety(abstract);分类 cs.LG

Comments To appear in the proceedings of the International Conference on Machine Learning and Cybernetics (ICMLC)

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.16410 2025-02-27 cs.CL 57%

Application of Multimodal Large Language Models in Autonomous Driving

Md Robiul Islam

机构 * William & Mary(威廉玛丽学院)

专题命中 其他安全 :safety(abstract);分类 cs.CL

Comments 9 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.19053 2025-02-27 physics.app-ph cs.ET cs.LG cs.NE 57%

Blending Optimal Control and Biologically Plausible Learning for Noise-Robust Physical Neural Networks

Satoshi Sunada, Tomoaki Niiyama, Kazutaka Kanno, Rin Nogami, André Röhm, Takato Awano, Atsushi Uchida

机构 * Kanazawa University(金泽大学) Saitama University(埼玉大学) The University of Tokyo(东京大学)

专题命中 其他安全 :alignment(abstract);分类 cs.LG

Comments 28 pages, 10 figures

Journal ref Phys. Rev. Lett., 134, 017301 (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.18844 2025-02-27 cs.CV cs.AI 57%

BarkXAI: A Lightweight Post-Hoc Explainable Method for Tree Species Classification with Quantifiable Concepts

Yunmei Huang, Songlin Hou, Zachary Nelson Horve, Songlin Fei

机构 * Purdue University(普渡大学) Worcester Polytechnic Institute(伍斯特理工学院)

专题命中 其他安全 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.18690 2025-02-27 cs.AI cs.RO 57%

Hybrid Voting-Based Task Assignment in Role-Playing Games

Daniel Weiner, Raj Korpan

机构 * Graduate Center, City University of New York(纽约城市大学研究生中心)

专题命中 其他安全 :alignment(abstract);分类 cs.AI

Comments Accepted for presentation at Dungeons, Neurons, and Dialogues: Social Interaction Dynamics in Contextual Games Workshop at 20th Annual ACM/IEEE International Conference on Human-Robot Interaction (HRI 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.05165 2025-02-27 cs.IR cs.CL 57%

Efficient Inference for Large Language Model-based Generative Recommendation

Xinyu Lin, Chaoqun Yang, Wenjie Wang, Yongqi Li, Cunxiao Du, Fuli Feng, See-Kiong Ng, Tat-Seng Chua

机构 * National University of Singapore(新加坡国立大学) Tsinghua University(清华大学) University of Science and Technology of China(中国科学技术大学) The Hong Kong Polytechnic University(香港理工大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL

Comments Accepted by ICLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.17883 2025-02-26 cs.CV cs.AI 57%

From underwater to aerial: a novel multi-scale knowledge distillation approach for coral reef monitoring

Matteo Contini, Victor Illien, Julien Barde, Sylvain Poulain, Serge Bernard, Alexis Joly, Sylvain Bonhommeau

机构 * IFREMER(法国海洋开发研究院) INRIA(法国国家信息与自动化研究所) Université de Montpellier(蒙彼利埃大学) CNRS(法国国家科学研究中心) IRD(法国发展研究院)

专题命中 其他安全 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.17505 2025-02-26 physics.ins-det cs.AI physics.data-an 57%

Inverse Surrogate Model of a Soft X-Ray Spectrometer using Domain Adaptation

Enrico Ahlers, Peter Feuer-Forson, Gregor Hartmann, Rolf Mitzner, Peter Baumgärtel, Jens Viefhaus

机构 * Helmholtz-Zentrum für Materialien und Energie GmbH(亥姆霍兹材料与能源中心) Humboldt-Universität zu Berlin(柏林洪堡大学)

专题命中 其他安全 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.13333 2025-02-26 cs.LG 57%

Beyond Accuracy: On the Effects of Fine-tuning Towards Vision-Language Model's Prediction Rationality

Qitong Wang, Tang Li, Kien X. Nguyen, Xi Peng

专题命中 其他安全 :safety(abstract);分类 cs.LG

Comments In Proceedings of the Association for the Advancement of Artificial Intelligence (AAAI), 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.12299 2025-02-26 cs.CL 57%

Semantics-Adaptive Activation Intervention for LLMs via Dynamic Steering Vectors

Weixuan Wang, Jingyuan Yang, Wei Peng

机构 * School of Informatics, University of Edinburgh(爱丁堡大学信息学院) Huawei Technologies Co., Ltd.(华为技术有限公司) School of Engineering, RMIT University(RMIT大学工程学院)

专题命中 其他安全 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.17239 2025-02-25 cs.CL cs.SD eess.AS 57%

Baichuan-Audio: A Unified Framework for End-to-End Speech Interaction

Tianpeng Li, Jun Liu, Tao Zhang, Yuanbo Fang, Da Pan, Mingrui Wang, Zheng Liang, Zehuan Li, Mingan Lin, Guosheng Dong, Jianhua Xu, Haoze Sun, Zenan Zhou, Weipeng Chen

机构 * Baichuan Inc.(百川智能)

专题命中 其他安全 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏