arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 9434 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 9434 篇

2503.00518 2025-03-04 cs.CV cs.LG 57%

Explainable LiDAR 3D Point Cloud Segmentation and Clustering for Detecting Airplane-Generated Wind Turbulence

Zhan Qu, Shuzhou Yuan, Michael Färber, Marius Brennfleck, Niklas Wartha, Anton Stephan

机构 * Karlsruhe Institute of Technology(卡尔斯鲁厄理工学院) ScaDS.AI, TU Dresden(德累斯顿工业大学ScaDS.AI) Institute of Atmospheric Physics, German Aerospace Center(德国航空航天中心大气物理研究所) RWTH Aachen University(亚琛工业大学)

专题命中 安全评测 :safety(abstract);分类 cs.LG

Comments Accepted at KDD 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.00384 2025-03-04 cs.CV cs.AI 57%

A Survey of Adversarial Defenses in Vision-based Systems: Categorization, Methods and Challenges

Nandish Chattopadhyay, Abdul Basit, Bassem Ouni, Muhammad Shafique

机构 * National Institute of Standards and Technology(美国国家标准与技术研究院) Colorado State University(科罗拉多州立大学)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.00330 2025-03-04 cs.CL 57%

How Deep is Love in LLMs' Hearts? Exploring Semantic Size in Human-like Cognition

Yao Yao, Yifei Yang, Xinbei Ma, Dongjie Yang, Zhuosheng Zhang, Zuchao Li, Hai Zhao

机构 * Shanghai Jiao Tong University(上海交通大学) Shanghai Key Laboratory of Trusted Data Circulation and Governance in Web3(上海Web3可信数据流通与治理重点实验室) Key Laboratory of Shanghai Education Commission for Intelligent Interaction and Cognitive Engineering(上海市教委智能交互与认知工程重点实验室)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.20975 2025-03-03 cs.CL 57%

Set-Theoretic Compositionality of Sentence Embeddings

Naman Bansal, Yash mahajan, Sanjeev Sinha, Santu Karmaker

机构 * Auburn University(奥本大学) University of Central Florida(中佛罗里达大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.20806 2025-03-03 cs.SE cs.AI 57%

Multimodal Learning for Just-In-Time Software Defect Prediction in Autonomous Driving Systems

Faisal Mohammad, Duksan Ryu

专题命中 安全评测 :safety(abstract);分类 cs.AI

Comments 9

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.18817 2025-02-27 cs.CL 57%

Judge as A Judge: Improving the Evaluation of Retrieval-Augmented Generation through the Judge-Consistency of Large Language Models

Shuliang Liu, Xinze Li, Zhenghao Liu, Yukun Yan, Cheng Yang, Zheni Zeng, Zhiyuan Liu, Maosong Sun, Ge Yu

机构 * Northeastern University(东北大学) Tsinghua University(清华大学) Beijing University of Posts and Telecommunications(北京邮电大学)

专题命中 安全评测 :DPO(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.18511 2025-02-27 cs.CR cs.AI 57%

ELBA-Bench: An Efficient Learning Backdoor Attacks Benchmark for Large Language Models

Xuxu Liu, Siyuan Liang, Mengya Han, Yong Luo, Aishan Liu, Xiantao Cai, Zheng He, Dacheng Tao

机构 * School of Computer Science, Wuhan University(武汉大学计算机学院) National University of Singapore(新加坡国立大学) Beihang University(北京航空航天大学) Nanyang Technological University(南洋理工大学)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.18228 2025-02-26 cs.CL 57%

Debt Collection Negotiations with Large Language Models: An Evaluation System and Optimizing Decision Making with Multi-Agent

Xiaofeng Wang, Zhixin Zhang, Jinguang Zheng, Yiming Ai, Rui Wang

机构 * Shanghai Jiao Tong University(上海交通大学) Ant Group(蚂蚁集团)

专题命中 安全评测 :DPO(abstract);分类 cs.CL

Comments 21 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.09179 2025-02-26 cs.LG 57%

Towards Effective Evaluations and Comparisons for LLM Unlearning Methods

Qizhou Wang, Bo Han, Puning Yang, Jianing Zhu, Tongliang Liu, Masashi Sugiyama

机构 * Hong Kong Baptist University(香港浸会大学) RIKEN Center for Advanced Intelligence Project(理化学研究所先进智能项目中心) The University of Sydney(悉尼大学) The University of Tokyo(东京大学)

专题命中 安全评测 :red teaming(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.05407 2025-02-26 eess.SP cs.CV cs.LG 57%

Machine Learning and Feature Ranking for Impact Fall Detection Event Using Multisensor Data

Tresor Y. Koffi, Youssef Mourchid, Mohammed Hindawi, Yohan Dupuis

机构 * CESI LINEACT Laboratory(CESI LINEACT实验室)

专题命中 安全评测 :safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.17646 2025-02-26 cs.LG cs.SE 57%

Architecting Digital Twins for Intelligent Transportation Systems

Hiya Bhatt, Sahil, Karthik Vaidhyanathan, Rahul Biju, Deepak Gangadharan, Ramona Trestian, Purav Shah

机构 * Software Engineering Research Centre, IIIT Hyderabad(印度信息技术与管理大学海得拉巴分校软件工程研究中心) Computer Systems Group, IIIT Hyderabad(印度信息技术与管理大学海得拉巴分校计算机系统组) Faculty of Science and Technology, Middlesex University London(伦敦密德萨斯大学科学与技术学院)

专题命中 安全评测 :safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.16402 2025-02-25 cs.AI 57%

Navigation-GPT: A Robust and Adaptive Framework Utilizing Large Language Models for Navigation Applications

Feng Ma, Xiu-min Wang, Chen Chen, Xiao-bin Xu, Xin-ping Yan

机构 * Wuhan University of Technology(武汉理工大学) Intelligent Transportation Systems Research Center, Wuhan University of Technology(武汉理工大学智能交通系统研究中心) Hangzhou Dianzi University(杭州电子科技大学) Nanjing Smart Water Transport Technology Co., Ltd.(南京智慧水运科技有限公司)

专题命中 安全评测 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.16064 2025-02-25 cs.LG 57%

Single Domain Generalization with Model-aware Parametric Batch-wise Mixup

Marzi Heidari, Yuhong Guo

机构 * School of Computer Science, Carleton University(卡尔顿大学计算机科学学院)

专题命中 安全评测 :alignment(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.14556 2025-02-25 cs.CL 57%

CitaLaw: Enhancing LLM with Citations in Legal Domain

Kepu Zhang, Weijie Yu, Sunhao Dai, Jun Xu

机构 * Gaoling School of Artificial Intelligence, Renmin University of China(中国人民大学高瓴人工智能学院) University of International Business and Economics(对外经济贸易大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.15545 2025-02-24 cs.CV cs.LG eess.IV 57%

Estimating Vehicle Speed on Roadways Using RNNs and Transformers: A Video-based Approach

Sai Krishna Reddy Mareddy, Dhanush Upplapati, Dhanush Kumar Antharam

机构 * University of North Carolina at Charlotte(北卡罗来纳大学夏洛特分校)

专题命中 安全评测 :safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.15427 2025-02-24 cs.CR cs.LG 57%

Adversarial Prompt Evaluation: Systematic Benchmarking of Guardrails Against Prompt Input Attacks on LLMs

Giulio Zizzo, Giandomenico Cornacchia, Kieran Fraser, Muhammad Zaid Hameed, Ambrish Rawat, Beat Buesser, Mark Purcell, Pin-Yu Chen, Prasanna Sattigeri, Kush Varshney

机构 * IBM Research(IBM研究院)

专题命中 安全评测 :jailbreak(abstract);分类 cs.LG

Comments NeurIPS 2024, Safe Generative AI Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.14437 2025-02-21 cs.CL 57%

Natural Language Generation

Ehud Reiter

专题命中 安全评测 :safety(abstract);分类 cs.CL

Comments This is a preprint of the following work: Ehud Reiter, Natural Language Generation, 2024, Springer reproduced with permission of Springer Nature Switzerland AG. The final authenticated version is available online at: http://dx.doi.org/10.1007/978-3-031-68582-8

Journal ref Book published by Springer in 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.13996 2025-02-21 cs.LG 57%

Beyond Single-Value Metrics: Evaluating and Enhancing LLM Unlearning with Cognitive Diagnosis

Yicheng Lang, Kehan Guo, Yue Huang, Yujun Zhou, Haomin Zhuang, Tianyu Yang, Yao Su, Xiangliang Zhang

机构 * University of Notre Dame(圣母大学) Worcester Polytechnic Institute(伍斯特理工学院)

专题命中 安全评测 :safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.07355 2025-02-21 cs.CL 57%

Think Together and Work Better: Combining Humans' and LLMs' Think-Aloud Outcomes for Effective Text Evaluation

SeongYeub Chu, JongWoo Kim, MunYong Yi

机构 * Graduate School of Data Science, KAIST(韩国科学技术院数据科学研究生院) Department of Industrial & Systems Engineering, KAIST(韩国科学技术院工业与系统工程系)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.13425 2025-02-21 cs.CV cs.AI 57%

Enhancing Adversarial Robustness of Vision-Language Models through Low-Rank Adaptation

Yuheng Ji, Yue Liu, Zhicheng Zhang, Zhao Zhang, Yuting Zhao, Xiaoshuai Hao, Gang Zhou, Xingwei Zhang, Xiaolong Zheng

机构 * Institute of Automation, CAS(中国科学院自动化研究所) School of Artificial Intelligence, UCAS(中国科学院大学人工智能学院) Institute of Data Science, National University of Singapore(新加坡国立大学数据科学研究所) Beijing University of Posts and Telecommunications(北京邮电大学) Beijing Academy of Artificial Intelligence(北京人工智能研究院)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.06656 2025-02-20 cs.AI 57%

A Frontier AI Risk Management Framework: Bridging the Gap Between Current AI Practices and Established Risk Management

Simeon Campos, Henry Papadatos, Fabien Roger, Chloé Touzet, Otter Quarks, Malcolm Murray

专题命中 安全评测 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.10009 2025-02-20 astro-ph.IM cs.AI cs.DL 57%

Enhancing Peer Review in Astronomy: A Machine Learning and Optimization Approach to Reviewer Assignments for ALMA

John M. Carpenter, Andrea Corvillón, Nihar B. Shah

机构 * Joint ALMA Observatory(联合ALMA天文台) Carnegie Mellon University(卡内基梅隆大学)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

Comments 16 pages, 5 figures, revised version accepted by PASP

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.13771 2025-02-20 cs.CV cs.LG 57%

Interpreting Neurons in Deep Vision Networks with Language Models

Nicholas Bai, Rahul A. Iyer, Tuomas Oikarinen, Akshay Kulkarni, Tsui-Wei Weng

机构 * UC San Diego(加州大学圣地亚哥分校) UT Austin(德克萨斯大学奥斯汀分校)

专题命中 安全评测 :trustworthy(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.12509 2025-02-19 cs.CL 57%

Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge

Kayla Schroeder, Zach Wood-Doughty

机构 * Northwestern University(西北大学)

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.12102 2025-02-18 cs.AI cs.ET 57%

Relational Norms for Human-AI Cooperation

Brian D. Earp, Sebastian Porsdam Mann, Mateo Aboy, Edmond Awad, Monika Betzler, Marietjie Botes, Rachel Calcott, Mina Caraccio, Nick Chater, Mark Coeckelbergh, Mihaela Constantinescu, Hossein Dabbagh, Kate Devlin, Xiaojun Ding, Vilius Dranseika, Jim A. C. Everett, Ruiping Fan, Faisal Feroz, Kathryn B. Francis, Cindy Friedman, Orsolya Friedrich, Iason Gabriel, Ivar Hannikainen, Julie Hellmann, Arasj Khodadade Jahrome, Niranjan S. Janardhanan, Paul Jurcys, Andreas Kappes, Maryam Ali Khan, Gordon Kraft-Todd, Maximilian Kroner Dale, Simon M. Laham, Benjamin Lange, Muriel Leuenberger, Jonathan Lewis, Peng Liu, David M. Lyreskog, Matthijs Maas, John McMillan, Emilian Mihailov, Timo Minssen, Joshua Teperowski Monrad, Kathryn Muyskens, Simon Myers, Sven Nyholm, Alexa M. Owen, Anna Puzio, Christopher Register, Madeline G. Reinecke, Adam Safron, Henry Shevlin, Hayate Shimizu, Peter V. Treit, Cristina Voinea, Karen Yan, Anda Zahiu, Renwen Zhang, Hazem Zohny, Walter Sinnott-Armstrong, Ilina Singh, Julian Savulescu, Margaret S. Clark

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

Comments 76 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.12073 2025-02-18 cs.CL 57%

Can LLMs Simulate Social Media Engagement? A Study on Action-Guided Response Generation

Zhongyi Qiu, Hanjia Lyu, Wei Xiong, Jiebo Luo

机构 * School of Computational Science and Engineering, Georgia Institute of Technology(佐治亚理工学院计算科学与工程学院) Department of Computer Science, University of Rochester(罗切斯特大学计算机科学系)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.11770 2025-02-18 cs.AI 57%

Cognitive-Aligned Document Selection for Retrieval-augmented Generation

Bingyu Wan, Fuxi Zhang, Zhongpeng Qi, Jiayi Ding, Jijun Li, Baoshi Fan, Yijia Zhang, Jun Zhang

机构 * Dalian Maritime University(大连海事大学)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.07191 2025-02-18 cs.AI 57%

Bag of Tricks for Inference-time Computation of LLM Reasoning

Fan Liu, Wenshuo Chao, Naiqiang Tan, Hao Liu

机构 * Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) Didichuxing Co. Ltd(滴滴出行有限公司)

专题命中 安全评测 :RLHF(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.12619 2025-02-18 cs.CL 57%

Quantification of Large Language Model Distillation

Sunbowen Lee, Junting Zhou, Chang Ao, Kaige Li, Xinrun Du, Sirui He, Haihong Wu, Tianci Liu, Jiaheng Liu, Hamid Alinejad-Rokny, Min Yang, Yitao Liang, Zhoufutu Wen, Shiwen Ni

机构 * Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences(中国科学院深圳先进技术研究院) Peking University(北京大学) SUSTech(南方科技大学) SUAT Leibowitz AI UNSW Sydney(新南威尔士大学悉尼分校)

专题命中 安全评测 :safety(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.10724 2025-02-18 cs.CL 57%

Large Language Models Are Active Critics in NLG Evaluation

Shuying Xu, Junjie Hu, Ming Jiang

机构 * Tongji University(同济大学) University of Wisconsin-Madison(威斯康星大学麦迪逊分校) Indiana University Indianapolis(印第安纳大学印第安纳波利斯分校)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏