arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 9434 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 9434 篇

2501.11213 2025-01-22 cs.LG 57%

Risk Analysis of Flowlines in the Oil and Gas Sector: A GIS and Machine Learning Approach

I. Chittumuri, N. Alshehab, R. J. Voss, L. L. Douglass, S. Kamrava, Y. Fan, J. Miskimins, W. Fleckenstein, S. Bandyopadhyay

机构 * Colorado School of Mines(科罗拉多矿业学院)

专题命中 安全评测 :safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.10897 2025-01-22 math.ST cs.LG stat.TH 57%

Unfolding Tensors to Identify the Graph in Discrete Latent Bipartite Graphical Models

Yuqi Gu

机构 * Columbia University(哥伦比亚大学)

专题命中 安全评测 :trustworthy(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.19740 2025-01-22 cs.CL 57%

Exploring the Reliability of Large Language Models as Customized Evaluators for Diverse NLP Tasks

Qintong Li, Leyang Cui, Lingpeng Kong, Wei Bi

机构 * The University of Hong Kong(香港大学) Tencent AI lab(腾讯AI实验室)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

Comments We release our resources at \url{https://github.com/qtli/CoEval}

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.09948 2025-01-20 eess.SY cs.AI cs.SY 57%

AI Explainability for Power Electronics: From a Lipschitz Continuity Perspective

Xinze Li, Fanfan Lin, Homer Alan Mantooth, Juan José Rodríguez-Andina

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.09355 2025-01-17 cs.AI cs.CV cs.ET cs.MA 57%

YETI (YET to Intervene) Proactive Interventions by Multimodal AI Agents in Augmented Reality Tasks

Saptarashmi Bandyopadhyay, Vikas Bahirwani, Lavisha Aggarwal, Bhanu Guda, Lin Li, Andrea Colaco

机构 * University of Maryland(马里兰大学) Google(谷歌公司)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

Comments Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.09294 2025-01-17 cs.CV cs.CL 57%

Efficient Few-Shot Medical Image Analysis via Hierarchical Contrastive Vision-Language Learning

Harrison Fuller, Fernando Gabriela Garcia, Victor Flores

机构 * Autonomous University of Nuevo León(新莱昂自治大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.09520 2025-01-17 cs.CV cs.AI 57%

Enhancing Skin Disease Diagnosis: Interpretable Visual Concept Discovery with SAM

Xin Hu, Janet Wang, Jihun Hamm, Rie R Yotsu, Zhengming Ding

机构 * Tulane University(杜兰大学)

专题命中 安全评测 :safety(abstract);分类 cs.AI

Comments This paper is accepted by WACV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.07769 2025-01-15 cs.LG cs.CV 57%

BMIP: Bi-directional Modality Interaction Prompt Learning for VLM

Song-Lin Lv, Yu-Yang Chen, Zhi Zhou, Ming Yang, Lan-Zhe Guo

机构 * School of Intelligence Science and Technology, Nanjing University(南京大学智能科学与技术学院) School of Artificial Intelligence, Nanjing University(南京大学人工智能学院) National Key Laboratory for Novel Software Technology, Nanjing University(南京大学计算机软件新技术国家重点实验室)

专题命中 安全评测 :alignment(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.08100 2025-01-15 quant-ph cs.LG 57%

Graph Neural Networks for Parameterized Quantum Circuits Expressibility Estimation

Shamminuj Aktar, Andreas Bärtschi, Diane Oyen, Stephan Eidenbenz, Abdel-Hameed A. Badawy

机构 * New Mexico State University(新墨西哥州立大学) Los Alamos National Laboratory(洛斯阿拉莫斯国家实验室)

专题命中 安全评测 :alignment(abstract);分类 cs.LG

Journal ref IEEE International Conference on Quantum Computing and Engineering (QCE), Montreal, QC, Canada, 2024, pp. 1547-1557

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.06907 2025-01-14 cs.LG 57%

Deep Learning and Foundation Models for Weather Prediction: A Survey

Jimeng Shi, Azam Shirali, Bowen Jin, Sizhe Zhou, Wei Hu, Rahuul Rangaraj, Shaowen Wang, Jiawei Han, Zhaonan Wang, Upmanu Lall, Yanzhao Wu, Leonardo Bobadilla, Giri Narasimhan

机构 * Florida International University(佛罗里达国际大学) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) New York University Shanghai(纽约大学上海分校) Arizona State University(亚利桑那州立大学) Columbia University(哥伦比亚大学)

专题命中 安全评测 :trustworthy(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.04911 2025-01-10 cs.CV cs.CY 57%

A Machine Learning Model for Crowd Density Classification in Hajj Video Frames

Afnan A. Shah

专题命中 安全评测 :safety(abstract);分类 cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.04811 2025-01-10 cs.LG cs.CR cs.SE 57%

Fast, Fine-Grained Equivalence Checking for Neural Decompilers

Luke Dramko, Claire Le Goues, Edward J. Schwartz

机构 * Carnegie Mellon University(卡内基梅隆大学) Software Engineering Institute(软件工程研究所)

专题命中 安全评测 :alignment(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.03402 2025-01-08 math.ST cs.LG stat.TH 57%

On the Adversarial Robustness of Benjamini Hochberg

Louis L Chen, Roberto Szechtman, Matan Seri

机构 * Naval Postgraduate School(海军研究生院)

专题命中 安全评测 :safety(abstract);分类 cs.LG

Comments 22 pages, 5 figures, NeurIPS 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.06538 2025-01-08 cs.LG cs.CR cs.CV 57%

Transferable Adversarial Examples with Bayes Approach

Mingyuan Fan, Cen Chen, Wenmeng Zhou, Yinggui Wang

机构 * East China Normal University(华东师范大学) Alibaba Group(阿里巴巴集团) Ant Group(蚂蚁集团)

专题命中 安全评测 :trustworthy(abstract);分类 cs.LG

Comments Accepted in AsiaCCS'25

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.01886 2025-01-06 cs.RO cs.AI cs.SY eess.SY 57%

Evaluating Scenario-based Decision-making for Interactive Autonomous Driving Using Rational Criteria: A Survey

Zhen Tian, Zhihao Lin, Dezong Zhao, Wenjing Zhao, David Flynn, Shuja Ansari, Chongfeng Wei

机构 * James Watt School of Engineering, University of Glasgow(格拉斯哥大学詹姆斯·瓦特工程学院)

专题命中 安全评测 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.00432 2025-01-03 cs.CV cs.LG 57%

OV-HHIR: Open Vocabulary Human Interaction Recognition Using Cross-modal Integration of Large Language Models

Lala Shakti Swarup Ray, Bo Zhou, Sungho Suh, Paul Lukowicz

机构 * RPTU Kaiserslautern-Landau(莱茵兰-普法尔茨州立大学凯撒斯劳滕-兰道分校) German Research Center for Artificial Intelligence (DFKI)(德国人工智能研究中心)

专题命中 安全评测 :safety(abstract);分类 cs.LG

Comments Accepted in IEEE ICASSP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.14285 2024-12-31 cs.RO cs.AI 57%

LLM-Personalize: Aligning LLM Planners with Human Preferences via Reinforced Self-Training for Housekeeping Robots

Dongge Han, Trevor McInroe, Adam Jelley, Stefano V. Albrecht, Peter Bell, Amos Storkey

专题命中 安全评测 :alignment(abstract);分类 cs.AI

Comments COLING 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.18727 2024-12-30 cs.SE cs.AI cs.SY eess.SY 57%

SAFLITE: Fuzzing Autonomous Systems via Large Language Models

Taohong Zhu, Adrians Skapars, Fardeen Mackenzie, Declan Kehoe, William Newton, Suzanne Embury, Youcheng Sun

专题命中 安全评测 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.18150 2024-12-30 cs.CV cs.AI 57%

EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation

Shuhao Han, Haotian Fan, Jiachen Fu, Liang Li, Tao Li, Junhui Cui, Yunqiu Wang, Yang Tai, Jingwei Sun, Chunle Guo, Chongyi Li

专题命中 安全评测 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.16469 2024-12-30 cs.CL 57%

Chained Tuning Leads to Biased Forgetting

Megan Ung, Alicia Sun, Samuel J. Bell, Bhaktipriya Radharapu, Levent Sagun, Adina Williams

专题命中 安全评测 :safety(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.18096 2024-12-25 cs.AI 57%

Real-world Deployment and Evaluation of PErioperative AI CHatbot (PEACH) -- a Large Language Model Chatbot for Perioperative Medicine

Yu He Ke, Liyuan Jin, Kabilan Elangovan, Bryan Wen Xi Ong, Chin Yang Oh, Jacqueline Sim, Kenny Wei-Tsen Loh, Chai Rick Soh, Jonathan Ming Hua Cheng, Aaron Kwang Yang Lee, Daniel Shu Wei Ting, Nan Liu, Hairil Rizal Abdullah

专题命中 安全评测 :safety(abstract);分类 cs.AI

Comments 21 pages, 3 figures, 1 graphical abstract

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.15446 2024-12-25 cs.CV cs.AI 57%

Concept Complement Bottleneck Model for Interpretable Medical Image Diagnosis

Hongmei Wang, Junlin Hou, Hao Chen

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

Comments 27 pages, 5 figures,

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.17787 2024-12-24 cs.CV cs.CL 57%

Cross-Lingual Text-Rich Visual Comprehension: An Information Theory Perspective

Xinmiao Yu, Xiaocheng Feng, Yun Li, Minghui Liao, Ya-Qi Yu, Xiachong Feng, Weihong Zhong, Ruihan Chen, Mengkang Hu, Jihao Wu, Dandan Tu, Duyu Tang, Bing Qin

专题命中 安全评测 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.17383 2024-12-24 cs.CL 57%

Interweaving Memories of a Siamese Large Language Model

Xin Song, Zhikai Xue, Guoxiu He, Jiawei Liu, Wei Lu

专题命中 安全评测 :alignment(abstract);分类 cs.CL

Comments Accepted by AAAI 2025 Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.17146 2024-12-24 cs.AI physics.flu-dyn 57%

LLM Agent for Fire Dynamics Simulations

Leidong Xu, Danyal Mohaddes, Yi Wang

专题命中 安全评测 :safety(abstract);分类 cs.AI

Comments NeurIPS 2024 Foundation Models for Science Workshop (38th Conference on Neural Information Processing Systems). 12 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.16762 2024-12-24 cs.RO cs.AI cs.SE 57%

A Method for the Runtime Validation of AI-based Environment Perception in Automated Driving System

Iqra Aslam, Abhishek Buragohain, Daniel Bamal, Adina Aniculaesei, Meng Zhang, Andreas Rausch

专题命中 安全评测 :safety(abstract);分类 cs.AI

Comments 9 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.16270 2024-12-24 cs.AI cs.HC 57%

MetaScientist: A Human-AI Synergistic Framework for Automated Mechanical Metamaterial Design

Jingyuan Qi, Zian Jia, Minqian Liu, Wangzhi Zhan, Junkai Zhang, Xiaofei Wen, Jingru Gan, Jianpeng Chen, Qin Liu, Mingyu Derek Ma, Bangzheng Li, Haohui Wang, Adithya Kulkarni, Muhao Chen, Dawei Zhou, Ling Li, Wei Wang, Lifu Huang

专题命中 安全评测 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.16243 2024-12-24 cs.LG 57%

Bag of Tricks for Multimodal AutoML with Image, Text, and Tabular Data

Zhiqiang Tang, Zihan Zhong, Tong He, Gerald Friedland

专题命中 安全评测 :alignment(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.16211 2024-12-24 cs.CV cs.CL cs.GR 57%

Is Your World Simulator a Good Story Presenter? A Consecutive Events-Based Benchmark for Future Long Video Generation

Yiping Wang, Xuehai He, Kuan Wang, Luyao Ma, Jianwei Yang, Shuohang Wang, Simon Shaolei Du, Yelong Shen

专题命中 安全评测 :alignment(abstract);分类 cs.CL

Comments benchmark paper, project page: https://ypwang61.github.io/project/StoryEval

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.15588 2024-12-23 cs.CL 57%

NeSyCoCo: A Neuro-Symbolic Concept Composer for Compositional Generalization

Danial Kamali, Elham J. Barezi, Parisa Kordjamshidi

专题命中 安全评测 :alignment(abstract);分类 cs.CL

Comments AAAI 2025 Project Page: https://iamdanialkamali.github.io/publication/neuro-symbolic-concept-composer

详情

展开后加载摘要…

URL PDF HTML 收藏