arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 9434 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 9434 篇

2506.14532 2025-06-18 cs.CL 57%

M2BeamLLM: Multimodal Sensing-empowered mmWave Beam Prediction with Large Language Models

Can Zheng, Jiguang He, Chung G. Kang, Guofa Cai, Zitong Yu, Merouane Debbah

机构 * Department of Electrical and Computer Engineering, Korea University(韩国大学电子与计算机工程系) School of Computing and Information Technology, Great Bay University(大湾大学计算机与信息科技学院) Dongguan Key Laboratory for Intelligence and Information Technology(东莞智能与信息科技重点实验室) Great Bay Institute for Advanced Study (GBIAS)(大湾先进研究学院) School of Information Engineering, Guangdong University of Technology(广东工业大学信息工程学院) Center for 6G Technology, Khalifa University of Science and Technology(哈里发科技大学6G技术中心)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

Comments 13 pages, 20 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.14074 2025-06-18 cs.LG cs.AR 57%

Comprehensive Verilog Design Problems: A Next-Generation Benchmark Dataset for Evaluating Large Language Models and Agents on RTL Design and Verification

Nathaniel Pinckney, Chenhui Deng, Chia-Tung Ho, Yun-Da Tsai, Mingjie Liu, Wenfei Zhou, Brucek Khailany, Haoxing Ren

机构 * NVIDIA

专题命中 安全评测 :alignment(abstract);分类 cs.LG

Comments 16 pages with appendix

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.13917 2025-06-18 cs.AI 57%

Evaluating Explainability: A Framework for Systematic Assessment and Reporting of Explainable AI Features

Miguel A. Lago, Ghada Zamzmi, Brandon Eich, Jana G. Delfino

机构 * U.S. Food and Drug Administration(美国食品药品监督管理局)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.11368 2025-06-18 cs.CL 57%

GuideBench: Benchmarking Domain-Oriented Guideline Following for LLM Agents

Lingxiao Diao, Xinyue Xu, Wanxuan Sun, Cheng Yang, Zhuosheng Zhang

机构 * School of Computer Science, Shanghai Jiao Tong University(上海交通大学计算机科学学院) Zhiyuan college, Shanghai Jiao Tong University(上海交通大学智远学院) ByteDance Inc(字节跳动公司)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

Comments ACL 2025 Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.13639 2025-06-17 cs.CL 57%

An Empirical Study of LLM-as-a-Judge: How Design Choices Impact Evaluation Reliability

Yusuke Yamauchi, Taro Yano, Masafumi Oyamada

机构 * The University of Tokyo(东京大学) NEC Corporation(日本电讯公司)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.13328 2025-06-17 cs.CL 57%

Document-Level Tabular Numerical Cross-Checking: A Coarse-to-Fine Approach

Chaoxu Pang, Yixuan Cao, Ganbin Zhou, Hongwei Li, Ping Luo

机构 * Key Laboratory of Intelligent Information Processing, Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所智能信息处理重点实验室) Beijing PAI Technology Ltd(北京PAI技术有限公司)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

Comments Submitted to IEEE TKDE

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.12486 2025-06-17 cs.AI 57%

DinoCompanion: An Attachment-Theory Informed Multimodal Robot for Emotionally Responsive Child-AI Interaction

Boyang Wang, Yuhao Song, Jinyuan Cao, Peng Yu, Hongcheng Guo, Zhoujun Li

机构 * Beihang University(北航) The University of Melbourne(墨尔本大学) Independent Researcher(独立研究者) Panasonic Appliances(China) Co.,Ltd(松下电器(中国)有限公司)

专题命中 安全评测 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.03847 2025-06-17 q-bio.QM cs.LG q-bio.BM 57%

Interpretable Multimodal Learning for Tumor Protein-Metal Binding: Progress, Challenges, and Perspectives

Xiaokun Liu, Sayedmohammadreza Rastegari, Yijun Huang, Sxe Chang Cheong, Weikang Liu, Wenjie Zhao, Qihao Tian, Hongming Wang, Yingjie Guo, Shuo Zhou, Sina Tabakhi, Xianyuan Liu, Zheqing Zhu, Wei Sang, Haiping Lu

机构 * Institute of Big Data Science and Industry, Shanxi University, Taiyuan, China(山西大学大数据科学与产业研究院) School of Computer and Information Technology, Shanxi University, Taiyuan, China(山西大学计算机与信息学院) Key Laboratory of Evolutionary Science Intelligence of Shanxi Province, Taiyuan, Shanxi, China(山西省进化科学智能重点实验室) Faculty of Computer Engineering, University of Isfahan, Isfahan, Iran(伊斯法罕大学计算机工程学院) The First Clinical Medical School, Shanxi Medical University, Taiyuan, China(山西医科大学第一临床医学院) School of Medicine & Population Health, University of Sheffield, Sheffield, UK(谢菲尔德大学医学院与人口健康学院) School of Computer Science, University of Sheffield, Sheffield, UK(谢菲尔德大学计算机科学学院) Centre for Machine Intelligence, University of Sheffield, Sheffield, UK(谢菲尔德大学智能中心)

专题命中 安全评测 :trustworthy(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.12152 2025-06-17 cs.AI 57%

Because we have LLMs, we Can and Should Pursue Agentic Interpretability

Been Kim, John Hewitt, Neel Nanda, Noah Fiedel, Oyvind Tafjord

专题命中 安全评测 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.12148 2025-06-17 cs.CL 57%

Hatevolution: What Static Benchmarks Don't Tell Us

Chiara Di Bonaventura, Barbara McGillivray, Yulan He, Albert Meroño-Peñuela

机构 * King’s College London(伦敦国王学院) Imperial College London(伦敦帝国学院)

专题命中 安全评测 :safety(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.07016 2025-06-17 cs.CV cs.AI 57%

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks

Sanjoy Chowdhury, Mohamed Elmoghany, Yohan Abeysinghe, Junjie Fei, Sayan Nag, Salman Khan, Mohamed Elhoseiny, Dinesh Manocha

机构 * University of Maryland, College Park(马里兰大学College Park分校) KAUST(卡尔斯兰大学) MBZUAI(马克斯·普朗克人工智能研究所) University of Toronto(多伦多大学)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

Comments Audio-visual learning, Audio-Visual RAG, Multi-Video Linkage

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.11957 2025-06-16 physics.med-ph cs.LG 57%

Automated Treatment Planning for Interstitial HDR Brachytherapy for Locally Advanced Cervical Cancer using Deep Reinforcement Learning

Mohammadamin Moradi, Runyu Jiang, Yingzi Liu, Malvern Madondo, Tianming Wu, James J. Sohn, Xiaofeng Yang, Yasmin Hasan, Zhen Tian

机构 * Department of Radiation & Cellular Oncology, University of Chicago(芝加哥大学辐射与细胞肿瘤学系) Department of Physics, University of Chicago(芝加哥大学物理系) Department of Radiation Oncology, Emory University(埃默里大学放射肿瘤学系)

专题命中 安全评测 :safety(abstract);分类 cs.LG

Comments 12 pages, 2 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.11763 2025-06-16 cs.CL cs.IR 57%

DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents

Mingxuan Du, Benfeng Xu, Chiwei Zhu, Xiaorui Wang, Zhendong Mao

专题命中 安全评测 :alignment(abstract);分类 cs.CL

Comments 31 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.11380 2025-06-16 cs.CV cs.AI 57%

Enhance Multimodal Consistency and Coherence for Text-Image Plan Generation

Xiaoxin Lu, Ranran Haoran Zhang, Yusen Zhang, Rui Zhang

机构 * The Pennsylvania State University(宾夕法尼亚州立大学)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

Comments 18 pages, 10 figures; Accepted to ACL 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.11020 2025-06-16 cs.SE cs.AI 57%

Extracting Knowledge Graphs from User Stories using LangChain

Thayná Camargo da Silva

专题命中 安全评测 :alignment(abstract);分类 cs.AI

Comments Master thesis work

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.10984 2025-06-16 cs.SE cs.AI 57%

Application Modernization with LLMs: Addressing Core Challenges in Reliability, Security, and Quality

Ahilan Ayyachamy Nadar Ponnusamy

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.10495 2025-06-16 cs.SE cs.AI 57%

How Well Do Large Language Models Serve as End-to-End Secure Code Agents for Python?

Jianian Gong, Nachuan Duan, Ziheng Tao, Zhaohui Gong, Yuan Yuan, Minlie Huang

机构 * School of Computer Science and Engineering, Beihang University(北京航空航天大学计算机科学与工程学院) State Key Laboratory of Software Development Environment(软件开发环境国家重点实验室) Zhongguancun Laboratory(中关村实验室) Tsinghua University(清华大学)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.08838 2025-06-13 cs.CL 57%

Disinformation Capabilities of Large Language Models

Ivan Vykopal, Matúš Pikuliak, Ivan Srba, Robert Moro, Dominik Macko, Maria Bielikova

机构 * Kempelen Institute of Intelligent Technologies(凯姆佩伦智能技术研究所) Faculty of Information Technology, Brno University of Technology(信息技术学院,布拉格技术大学)

专题命中 安全评测 :safety(abstract);分类 cs.CL

Journal ref Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.10282 2025-06-13 cs.LG 57%

Graph-MLLM: Harnessing Multimodal Large Language Models for Multimodal Graph Learning

Jiajin Liu, Dongzhe Fan, Jiacheng Shen, Chuanhao Ji, Daochen Zha, Qiaoyu Tan

机构 * New York University Shanghai(纽约大学上海校区) University of Chinese Academy of Sciences(中国科学院大学) Rice University(德克萨斯大学里德学院)

专题命中 安全评测 :alignment(abstract);分类 cs.LG

Comments 16 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.06971 2025-06-13 cs.CL cs.CR 57%

Chain-of-Code Collapse: Reasoning Failures in LLMs via Adversarial Prompting in Code Generation

Jaechul Roh, Varun Gandhi, Shivani Anilkumar, Arin Garg

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.17604 2025-06-13 cs.LG 57%

RmGPT: A Foundation Model with Generative Pre-trained Transformer for Fault Diagnosis and Prognosis in Rotating Machinery

Yilin Wang, Yifei Yu, Kong Sun, Peixuan Lei, Yuxuan Zhang, Enrico Zio, Aiguo Xia, Yuanxiang Li

机构 * School of Aeronautics and Astronautics, Shanghai Jiao Tong University(航空宇航学院,上海交通大学) Shanghai Innovation Institute(上海创新研究院) MINES Paris-PSL University(巴黎-里昂商学院) Energy Department, Politecnico di Milano(能源部门,米兰理工学院) Beijing Aeronautical Technology Research Center(北京航空航天科技研究所)

专题命中 安全评测 :safety(abstract);分类 cs.LG

Comments This paper has been accepted for publication in the IEEE Internet of Things Journal (IoT-J). The final version may differ slightly due to editorial revisions. Please cite the journal version when available

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.08174 2025-06-12 cs.CL 57%

LLM-BT-Terms: Back-Translation as a Framework for Terminology Standardization and Dynamic Semantic Embedding

Li Weigang, Pedro Carvalho Brom

机构 * TransLab, Computer Science Department University of Brasilia(翻译实验室、计算机科学系巴西大学) Math Department Federal Institute of Brasilia(数学系巴西联邦理工学院)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

Comments 23 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.21025 2025-06-12 cs.CL 57%

Durghotona GPT: A Web Scraping and Large Language Model Based Framework to Generate Road Accident Dataset Automatically in Bangladesh

MD Thamed Bin Zaman Chowdhury, Moazzem Hossain, Md. Ridwanul Islam

专题命中 安全评测 :safety(abstract);分类 cs.CL

Comments It has been accepted in IEEE 27th International Conference on Computer and Information Technology (ICCIT). Now, we are waiting for it to get published in IEEE Xplore

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.09414 2025-06-12 cs.CL cs.IR 57%

PGDA-KGQA: A Prompt-Guided Generative Framework with Multiple Data Augmentation Strategies for Knowledge Graph Question Answering

Xiujun Zhou, Pingjian Zhang, Deyou Tang

机构 * School of Software Engineering(软件工程学院) South China University of Technology(华南理工大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

Comments 13 pages, 7 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.08189 2025-06-11 cs.CV cs.CL 57%

Open World Scene Graph Generation using Vision Language Models

Amartya Dutta, Kazi Sajeed Mehrab, Medha Sawhney, Abhilash Neog, Mridul Khurana, Sepideh Fatemi, Aanish Pradhan, M. Maruf, Ismini Lourentzou, Arka Daw, Anuj Karpatne

专题命中 安全评测 :alignment(abstract);分类 cs.CL

Comments Accepted in CVPR 2025 Workshop (CVinW)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.07947 2025-06-10 cs.CL 57%

Statistical Hypothesis Testing for Auditing Robustness in Language Models

Paulius Rauba, Qiyao Wei, Mihaela van der Schaar

机构 * University of Cambridge(剑桥大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

Comments arXiv admin note: substantial text overlap with arXiv:2412.00868

Journal ref Forty-second International Conference on Machine Learning. ICML 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.07645 2025-06-10 cs.CL 57%

Evaluating LLMs Robustness in Less Resourced Languages with Proxy Models

Maciej Chrabąszcz, Katarzyna Lorenc, Karolina Seweryn

专题命中 安全评测 :safety(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.07631 2025-06-10 cs.CL cs.CV 57%

Unblocking Fine-Grained Evaluation of Detailed Captions: An Explaining AutoRater and Critic-and-Revise Pipeline

Brian Gordon, Yonatan Bitton, Andreea Marzoca, Yasumasa Onoe, Xiao Wang, Daniel Cohen-Or, Idan Szpektor

机构 * Tel Aviv University(特拉维夫大学) Google Research(谷歌研究)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.16789 2025-06-10 cs.CE cs.AI 57%

AlphaAgent: LLM-Driven Alpha Mining with Regularized Exploration to Counteract Alpha Decay

Ziyi Tang, Zechuan Chen, Jiarui Yang, Jiayao Mai, Yongsen Zheng, Keze Wang, Jinrui Chen, Liang Lin

机构 * Sun Yat-sen University(中山大学) University of New South Wales(新南威尔士大学) Nanyang Technological University(南洋理工大学) The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))

专题命中 安全评测 :alignment(abstract);分类 cs.AI

Comments 9 pages; Code is available at: https://github.com/RndmVariableQ/AlphaAgent

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.06830 2025-06-10 cs.CV cs.AI 57%

EndoARSS: Adapting Spatially-Aware Foundation Model for Efficient Activity Recognition and Semantic Segmentation in Endoscopic Surgery

Guankun Wang, Rui Tang, Mengya Xu, Long Bai, Huxin Gao, Hongliang Ren

机构 * Department of Electronic Engineering, The Chinese University of Hong Kong(香港中文大学电子工程系) Shenzhen Research Institute, The Chinese University of Hong Kong(香港中文大学深圳研究院) School of Advanced Manufacturing, Fuzhou University(福州大学先进制造学院)

专题命中 安全评测 :safety(abstract);分类 cs.AI

Comments Accepted by Advanced Intelligent Systems

详情

展开后加载摘要…

URL PDF HTML 收藏