arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 9434 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 9434 篇

2503.22925 2025-04-07 cs.RO cs.AI 57%

Predictive Traffic Rule Compliance using Reinforcement Learning

Yanliang Huang, Sebastian Mair, Zhuoqi Zeng, Matthias Althoff

专题命中 安全评测 :safety(abstract);分类 cs.AI

Comments 12 pages, 7 figures. Preprint intended for submission to IEEE ITSC 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.02733 2025-04-04 cs.CL 57%

Enhancing LLM Robustness to Perturbed Instructions: An Empirical Study

Aryan Agrawal, Lisa Alazraki, Shahin Honarvar, Marek Rei

机构 * Imperial College London(伦敦帝国学院)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

Comments Building Trust Workshop, ICLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.01349 2025-04-03 cs.CL 57%

Tasks and Roles in Legal AI: Data Curation, Annotation, and Verification

Allison Koenecke, Jed Stiglitz, David Mimno, Matthew Wilkens

机构 * Cornell University(康奈尔大学) Law School, Cornell University(康奈尔大学法学院) Department of Information Science, Cornell University(康奈尔大学信息科学系)

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.08678 2025-04-02 cs.LG math.OC stat.CO stat.ML 57%

How to beat a Bayesian adversary

Zihan Ding, Kexin Jin, Jonas Latz, Chenguang Liu

机构 * Princeton University(普林斯顿大学) University of Manchester(曼彻斯特大学) Technische Universiteit Delft(代尔夫特理工大学)

专题命中 安全评测 :safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.00018 2025-04-02 cs.CR cs.LG 57%

SandboxEval: Towards Securing Test Environment for Untrusted Code

Rafiqul Rabin, Jesse Hostetler, Sean McGregor, Brett Weir, Nick Judd

机构 * Digital Safety Research Institute(数字安全研究所) UL Research Institutes(UL研究机构)

专题命中 安全评测 :safety(abstract);分类 cs.LG

Comments preliminary version, working paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.23046 2025-04-01 cs.RO cs.LG 57%

VLM-C4L: Continual Core Dataset Learning with Corner Case Optimization via Vision-Language Models for Autonomous Driving

Haibo Hu, Jiacheng Zuo, Yang Lou, Yufei Cui, Jianping Wang, Nan Guan, Jin Wang, Yung-Hui Li, Chun Jason Xue

机构 * City University of Hong Kong(香港城市大学) Soochow University(苏州大学) McGill University(麦吉尔大学) Hon Hai Research Institute(鸿海研究院) Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)

专题命中 安全评测 :safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.22731 2025-04-01 cs.LG 57%

MoRE-LLM: Mixture of Rule Experts Guided by a Large Language Model

Alexander Koebler, Ingo Thon, Florian Buettner

机构 * Goethe University Frankfurt(法兰克福大学) Siemens AG(西门子公司) German Cancer Research Center (DKFZ)(德国癌症研究中心)

专题命中 安全评测 :alignment(abstract);分类 cs.LG

Comments 2024 IEEE International Conference on Data Mining (ICDM)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.22689 2025-04-01 cs.LG physics.data-an stat.AP 57%

From Occurrence to Consequence: A Comprehensive Data-driven Analysis of Building Fire Risk

Chenzhi Ma, Hongru Du, Shengzhi Luan, Ensheng Dong, Lauren M. Gardner, Thomas Gernay

机构 * Johns Hopkins University(约翰斯·霍普金斯大学) Boston University(波士顿大学) RAND Corporation(兰德公司) Johns Hopkins Bloomberg School of Public Health(约翰斯·霍普金斯大学布隆伯格公共卫生学院)

专题命中 安全评测 :safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.11145 2025-04-01 cs.CV cs.AI cs.MM 57%

Detecting Multimodal Situations with Insufficient Context and Abstaining from Baseless Predictions

Junzhang Liu, Zhecan Wang, Hammad Ayyubi, Haoxuan You, Chris Thomas, Rui Sun, Shih-Fu Chang, Kai-Wei Chang

机构 * Columbia University(哥伦比亚大学) University of California, Los Angeles(加利福尼亚大学洛杉矶分校)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.11498 2025-04-01 cs.RO cs.AI 57%

Verifiably Following Complex Robot Instructions with Foundation Models

Benedict Quartey, Eric Rosen, Stefanie Tellex, George Konidaris

机构 * Brown University(布朗大学)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.12523 2025-03-31 cs.CV cs.AI 57%

Single Image Unlearning: Efficient Machine Unlearning in Multimodal Large Language Models

Jiaqi Li, Qianshan Wei, Chuanyi Zhang, Guilin Qi, Miaozeng Du, Yongrui Chen, Sheng Bi, Fan Liu

专题命中 安全评测 :jailbreak(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.21613 2025-03-28 cs.CL 57%

Evaluating book summaries from internal knowledge in Large Language Models: a cross-model and semantic consistency approach

Javier Coronado-Blázquez

专题命中 安全评测 :alignment(abstract);分类 cs.CL

Comments 22 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.11441 2025-03-28 cs.IR cs.CL 57%

Ontology Matching with Large Language Models and Prioritized Depth-First Search

Maria Taboada, Diego Martinez, Mohammed Arideh, Rosa Mosquera

机构 * University of Santiago de Compostela(圣地亚哥德孔波斯特拉大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.18227 2025-03-27 cs.CV cs.AI 57%

PG-SAM: Prior-Guided SAM with Medical for Multi-organ Segmentation

Yiheng Zhong, Zihong Luo, Chengzhi Liu, Feilong Tang, Zelin Peng, Ming Hu, Yingzhen Hu, Jionglong Su, Zongyuan Ge, Imran Razzak

机构 * Mohamed bin Zayed University of AI(穆罕默德·本·扎耶德人工智能大学) Xi’an Jiaotong-Liverpool University(西交利物浦大学) University of Liverpool(利物浦大学) Monash University(莫纳什大学) Shanghai Jiao Tong University(上海交通大学)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.10013 2025-03-27 cs.CL 57%

Probabilistic Lexical Manifold Construction in Large Language Models via Hierarchical Vector Field Interpolation

Clive Pendleton, Ewan Harrington, Giles Fairbrother, Jasper Arkwright, Nigel Fenwick, Richard Katrix

专题命中 安全评测 :alignment(abstract);分类 cs.CL

Comments arXiv admin note: This paper has been withdrawn by arXiv due to disputed and unverifiable authorship

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.19478 2025-03-26 cs.CV cs.CY 57%

TeLL Me what you cant see

Saverio Cavasin, Pietro Biasetton, Mattia Tamiazzo, Mauro Conti, Simone Milani

机构 * University of Padua(帕多瓦大学)

专题命中 安全评测 :trustworthy(abstract);分类 cs.CY

Comments 16 pages, 58 images

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.18958 2025-03-26 cs.AI math.PR stat.ML 57%

Advancing Deep Learning through Probability Engineering: A Pragmatic Paradigm for Modern AI

Jianyi Zhang

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

Comments Ph.D. dissertation

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.16861 2025-03-26 cs.AI 57%

In-House Evaluation Is Not Enough: Towards Robust Third-Party Flaw Disclosure for General-Purpose AI

Shayne Longpre, Kevin Klyman, Ruth E. Appel, Sayash Kapoor, Rishi Bommasani, Michelle Sahar, Sean McGregor, Avijit Ghosh, Borhane Blili-Hamelin, Nathan Butters, Alondra Nelson, Amit Elazari, Andrew Sellars, Casey John Ellis, Dane Sherrets, Dawn Song, Harley Geiger, Ilona Cohen, Lauren McIlvenny, Madhulika Srikumar, Mark M. Jaycox, Markus Anderljung, Nadine Farid Johnson, Nicholas Carlini, Nicolas Miailhe, Nik Marda, Peter Henderson, Rebecca S. Portnoff, Rebecca Weiss, Victoria Westerhoff, Yacine Jernite, Rumman Chowdhury, Percy Liang, Arvind Narayanan

专题命中 安全评测 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.05395 2025-03-26 cs.CL 57%

Hierarchical Lexical Manifold Projection in Large Language Models: A Novel Mechanism for Multi-Scale Semantic Representation

Natasha Martus, Sebastian Crowther, Maxwell Dorrington, Jonathan Applethwaite, Edgar Tillinghurst, Quentin Birkenshaw, Lukas Petrov, Constance Willoughby

专题命中 安全评测 :alignment(abstract);分类 cs.CL

Comments arXiv admin note: This paper has been withdrawn by arXiv due to disputed and unverifiable authorship

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.05761 2025-03-26 cs.CL 57%

The BiGGen Bench: A Principled Benchmark for Fine-grained Evaluation of Language Models with Language Models

Seungone Kim, Juyoung Suk, Ji Yong Cho, Shayne Longpre, Chaeeun Kim, Dongkeun Yoon, Guijin Son, Yejin Cho, Sheikh Shafayat, Jinheon Baek, Sue Hyun Park, Hyeonbin Hwang, Jinkyung Jo, Hyowon Cho, Haebin Shin, Seongyun Lee, Hanseok Oh, Noah Lee, Namgyu Ho, Se June Joo, Miyoung Ko, Yoonjoo Lee, Hyungjoo Chae, Jamin Shin, Joel Jang, Seonghyeon Ye, Bill Yuchen Lin, Sean Welleck, Graham Neubig, Moontae Lee, Kyungjae Lee, Minjoon Seo

专题命中 安全评测 :harmlessness(abstract);分类 cs.CL

Comments NAACL 2025 (Main Conference)

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.01697 2025-03-26 cs.LG cs.SE cs.SY eess.SY 57%

Exploring Robustness of Image Recognition Models on Hardware Accelerators

Nikolaos Louloudakis, Perry Gibson, José Cano, Ajitha Rajan

机构 * University of Edinburgh(爱丁堡大学) University of Glasgow(格拉斯哥大学)

专题命中 安全评测 :safety(abstract);分类 cs.LG

Comments 7 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.21321 2025-03-25 cs.CL cs.CV 57%

LLM Post-Training: A Deep Dive into Reasoning Large Language Models

Komal Kumar, Tajamul Ashraf, Omkar Thawakar, Rao Muhammad Anwer, Hisham Cholakkal, Mubarak Shah, Ming-Hsuan Yang, Phillip H. S. Torr, Fahad Shahbaz Khan, Salman Khan

机构 * Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学) University of Central Florida(中佛罗里达大学) University of California at Merced(加州大学默塞德分校) Google DeepMind(谷歌DeepMind) University of Oxford(牛津大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

Comments 32 pages, 7 figures, 3 tables, 377 references. Github Repo: https://github.com/mbzuai-oryx/Awesome-LLM-Post-training

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.03871 2025-03-24 cs.CV cs.AI cs.ET cs.IR cs.MM 57%

CLIP-PING: Boosting Lightweight Vision-Language Models with Proximus Intrinsic Neighbors Guidance

Chu Myaet Thwal, Ye Lin Tun, Minh N. H. Nguyen, Eui-Nam Huh, Choong Seon Hong

机构 * Kyung Hee University(庆熙大学) The University of Danang—Vietnam-Korea University of Information and Communication Technology(岘港大学—越韩信息与通信技术大学) Digital Science and Technology Institute(数字科学技术研究院)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

Comments 14 pages, 5 figures, 24 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.09818 2025-03-24 cs.CL 57%

Chameleon: Mixed-Modal Early-Fusion Foundation Models

Chameleon Team

机构 * FAIR at Meta(Meta FAIR研究院)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.15793 2025-03-24 cs.HC cs.AI cs.IR 57%

Summaries, Highlights, and Action items: Design, implementation and evaluation of an LLM-powered meeting recap system

Sumit Asthana, Sagih Hilleli, Pengcheng He, Aaron Halfaker

机构 * University of Michigan(密歇根大学) Microsoft(微软公司)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

Comments Accepted at CSCW 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.15528 2025-03-21 cs.HC cs.AI 57%

Complying with the EU AI Act: Innovations in Explainable and User-Centric Hand Gesture Recognition

Sarah Seifi, Tobias Sukianto, Cecilia Carbonelli, Lorenzo Servadei, Robert Wille

机构 * Technical University Munich(慕尼黑工业大学) Infineon Technologies AG(英飞凌科技股份公司) Johannes Kepler University Linz(约翰内斯·开普勒大学林茨分校) Software Competence Center Hagenberg GmbH (SCCH)(哈根堡软件能力中心有限公司(SCCH))

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.15248 2025-03-20 cs.SE cs.AI 57%

Automated Non-Functional Requirements Generation in Software Engineering with Large Language Models: A Comparative Study

Jomar Thomas Almonte, Santhosh Anitha Boominathan, Nathalia Nascimento

机构 * The Pennsylvania State University(宾夕法尼亚州立大学)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

Comments 11 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.14797 2025-03-20 cs.CL 57%

FACTS&EVIDENCE: An Interactive Tool for Transparent Fine-Grained Factual Verification of Machine-Generated Text

Varich Boonsanong, Vidhisha Balachandran, Xiaochuang Han, Shangbin Feng, Lucy Lu Wang, Yulia Tsvetkov

机构 * University of Washington(华盛顿大学) Microsoft Research(微软研究院)

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.19634 2025-03-20 cs.CV cs.AI 57%

MedVLM-R1: Incentivizing Medical Reasoning Capability of Vision-Language Models (VLMs) via Reinforcement Learning

Jiazhen Pan, Che Liu, Junde Wu, Fenglin Liu, Jiayuan Zhu, Hongwei Bran Li, Chen Chen, Cheng Ouyang, Daniel Rueckert

机构 * Technical University of Munich (TUM)(慕尼黑工业大学(TUM)) TUM University Hospital(慕尼黑工业大学医院) University of Oxford(牛津大学) Imperial College London(伦敦帝国学院) Massachusetts General Hospital(麻省总医院) Harvard Medical School(哈佛医学院) University of Sheffield(谢菲尔德大学)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.03113 2025-03-19 cs.LG 57%

Predicting Space Tourism Demand Using Explainable AI

Tan-Hanh Pham, Jingchen Bi, Rodrigo Mesa-Arango, Kim-Doang Nguyen

机构 * Florida Institute of Technology(佛罗里达理工学院)

专题命中 安全评测 :trustworthy(abstract);分类 cs.LG

Comments 15 pages

详情

展开后加载摘要…

URL PDF HTML 收藏