arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 9346 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 9346 篇

2507.17216 2025-07-24 cs.CL cs.AI 62%

The Pluralistic Moral Gap: Understanding Judgment and Value Differences between Humans and Large Language Models

Giuseppe Russo, Debora Nozza, Paul Röttger, Dirk Hovy

机构 * EPFL(苏黎世联邦理工学院) Bocconi University(博科尼大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

Comments 13 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.05119 2025-07-23 cs.LG cs.AI cs.AR cs.CV eess.IV 62%

Balancing Robustness and Efficiency in Embedded DNNs Through Activation Function Selection

Jon Gutiérrez-Zaballa, Koldo Basterretxea, Javier Echanobe

机构 * Department of Electronics Technology, University of the Basque Country (UPV/EHU)(电子技术系,巴斯克国家大学(UPV/EHU)) Department of Electricity and Electronics, University of the Basque Country (UPV/EHU)(电力与电子系,巴斯克国家大学(UPV/EHU))

专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.15874 2025-07-23 cs.AI cs.CL 62%

Why Braking? Scenario Extraction and Reasoning Utilizing LLM

Yin Wu, Daniel Slieter, Vivek Subramanian, Ahmed Abouelazm, Robin Bohn, J. Marius Zöllner

机构 * CARIAD SE(CARIAD公司) Karlsruhe Institute of Technology(卡尔斯鲁厄理工学院) FZI Research Center for Information Technology(弗劳恩霍夫信息技术研究中心)

专题命中 安全评测 :safety(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.15868 2025-07-23 cs.CL cs.AI 62%

Small Edits, Big Consequences: Telling Good from Bad Robustness in Large Language Models

Altynbek Ismailov, Salia Asanova

机构 * Berkeley(伯克利)

专题命中 安全评测 :safety(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.15255 2025-07-22 eess.SP cs.AI cs.LG 62%

MEETI: A Multimodal ECG Dataset from MIMIC-IV-ECG with Signals, Images, Features and Interpretations

Deyun Zhang, Xiang Lan, Shijia Geng, Qinghao Zhao, Sumei Fan, Mengling Feng, Shenda Hong

机构 * HeartVoice Medical Technology(HeartVoice医疗科技) Saw Swee Hock School of Public Health and Institute of Data Science(Saw Swee Hock公共卫生学院和数据科学研究所) National University of Singapore(新加坡国立大学) Department of Cardiology, Peking University People’s Hospital(北京大学人民医院心内科) College of Integrative Chinese and Western Medicine, Anhui University of Chinese Medicine(安徽中医药大学整合中西医学学院) National Institute of Health Data Science, Peking University(北京大学国家健康数据科学研究院) Institute for Artificial Intelligence, Peking University(北京大学人工智能研究院)

专题命中 安全评测 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14824 2025-07-22 cs.LG cs.AI 62%

Benchmarking Foundation Models with Multimodal Public Electronic Health Records

Kunyu Yu, Rui Yang, Jingchi Liao, Siqi Li, Huitao Li, Irene Li, Yifan Peng, Rishikesan Kamaleswaran, Nan Liu

机构 * Centre for Quantitative Medicine and Duke-NUS AI + Medical Science Initiative, Duke-NUS Medical School(定量医学中心和杜克-国立新加坡大学AI+医学科学计划,杜克-国立新加坡大学医学院) Graduate School of Engineering, The University of Tokyo(东京大学工程研究生院) Department of Population Health Sciences, Weill Cornell Medicine(流行病学与公共卫生科学系,韦尔·科恩医学中心) Department of Surgery, Duke University School of Medicine(外科医学系,杜克大学医学学院) Centre for Quantitative Medicine, Duke-NUS AI + Medical Science Initiative and Programme in Health Services and Systems Research, Duke-NUS Medical School and NUS Artificial Intelligence Institute, National University of Singapore(定量医学中心和杜克-国立新加坡大学AI+医学科学计划及健康服务与系统研究计划,杜克-国立新加坡大学医学院和新加坡国立大学人工智能研究所)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14298 2025-07-22 cs.CL cs.AI cs.CV 62%

In-Depth and In-Breadth: Pre-training Multimodal Language Models Customized for Comprehensive Chart Understanding

Wan-Cyuan Fan, Yen-Chun Chen, Mengchen Liu, Alexander Jacobson, Lu Yuan, Leonid Sigal

机构 * UBC(不列颠哥伦比亚大学) Microsoft(微软) Vector Institute for AI(人工智能向量研究所) CIFAR AI Chair(卡尔·弗雷德里克人工智能主席)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

Comments arXiv admin note: substantial text overlap with arXiv:2407.14506

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14180 2025-07-22 cs.LG cs.AI 62%

Digital Twin-Assisted Explainable AI for Robust Beam Prediction in mmWave MIMO Systems

Nasir Khan, Asmaa Abdallah, Abdulkadir Celik, Ahmed M. Eltawil, Sinem Coleri

机构 * department of Electrical and Electronics Engineering, Koc University(电子与电气工程系,科克大学) Computer, Electrical, and Mathematical Sciences and Engineering Division, King Abdullah University of Science and Technology(计算机、电气和数学科学与工程系,国王阿卜杜勒-阿齐兹大学)

专题命中 安全评测 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.19027 2025-07-22 cs.AI cs.LG cs.NE 62%

DiCE-Extended: A Robust Approach to Counterfactual Explanations in Machine Learning

Volkan Bakir, Polat Goktas, Sureyya Akyuz

机构 * Faculty of Graduate Education Institute, Department of Artificial Intelligence (Interdisciplinary), Bahçeşehir University, Turkey(研究生教育学院人工智能系(跨学科)巴塞希尔大学,土耳其) School of Computer Science, University College Dublin, Ireland(计算机科学学院都柏林大学学院,爱尔兰) Faculty of Engineering and Natural Sciences, Department of Mathematics, Bahçeşehir University, Turkey(工程与自然科学学院数学系巴塞希尔大学,土耳其)

专题命中 安全评测 :alignment(abstract);分类 cs.AI、cs.LG

Comments 5th international Conference on Modelling, Computation and Optimization in Information Systems and Management Sciences (MCO 2025), June 4-6, 2025, Metz, France

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.08208 2025-07-21 cs.CL cs.AI 62%

ASTRID -- An Automated and Scalable TRIaD for the Evaluation of RAG-based Clinical Question Answering Systems

Mohita Chowdhury, Yajie Vera He, Jared Joselowitz, Aisling Higham, Ernest Lim

机构 * Ufonia Limited(乌菲尼亚有限公司) University of York(约克大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

Comments 29 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.13390 2025-07-21 cs.CL cs.LG 62%

PARAM-1 BharatGen 2.9B Model

Kundeshwar Pundalik, Piyush Sawarkar, Nihar Sahoo, Abhishek Shinde, Prateek Chanda, Vedant Goswami, Ajay Nagpal, Atul Singh, Viraj Thakur, Vijay Dewane, Aamod Thakur, Bhargav Patel, Smita Gautam, Bhagwan Panditi, Shyam Pawar, Madhav Kotcha, Suraj Racha, Saral Sureka, Pankaj Singh, Rishi Bal, Rohit Saluja, Ganesh Ramakrishnan

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.12674 2025-07-21 cs.CY cs.AI cs.SE 62%

ParaStudent: Generating and Evaluating Realistic Student Code by Teaching LLMs to Struggle

Mihran Miroyan, Rose Niousha, Joseph E. Gonzalez, Gireeja Ranade, Narges Norouzi

机构 * University of California, Berkeley(加州大学伯克利分校)

专题命中 安全评测 :alignment(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.14506 2025-07-21 cs.CV cs.AI cs.CL 62%

On Pre-training of Multimodal Language Models Customized for Chart Understanding

Wan-Cyuan Fan, Yen-Chun Chen, Mengchen Liu, Lu Yuan, Leonid Sigal

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

Comments NeurIPS 2024 Workshop on Adaptive Foundation Models

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.13090 2025-07-18 cs.LG cs.AI cs.CV 62%

MUPAX: Multidimensional Problem Agnostic eXplainable AI

Vincenzo Dentamaro, Felice Franchini, Giuseppe Pirlo, Irina Voiculescu

机构 * University of Oxford(牛津大学)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.19982 2025-07-17 cs.CL cs.AI 62%

TD-EVAL: Revisiting Task-Oriented Dialogue Evaluation by Combining Turn-Level Precision with Dialogue-Level Comparisons

Emre Can Acikgoz, Carl Guo, Suvodip Dey, Akul Datta, Takyoung Kim, Gokhan Tur, Dilek Hakkani-Tür

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.10502 2025-07-17 cs.LG cs.AI 62%

Benchmarking and Evaluation of AI Models in Biology: Outcomes and Recommendations from the CZI Virtual Cells Workshop

Elizabeth Fahsbender, Alma Andersson, Jeremy Ash, Polina Binder, Daniel Burkhardt, Benjamin Chang, Georg K. Gerber, Anthony Gitter, Patrick Godau, Ankit Gupta, Genevieve Haliburton, Siyu He, Trey Ideker, Ivana Jelic, Aly Khan, Yang-Joon Kim, Aditi Krishnapriyan, Jon M. Laurent, Tianyu Liu, Emma Lundberg, Shalin B. Mehta, Rob Moccia, Angela Oliveira Pisco, Katherine S. Pollard, Suresh Ramani, Julio Saez-Rodriguez, Yasin Senbabaoglu, Elana Simon, Srinivasan Sivanandan, Gustavo Stolovitzky, Marc Valer, Bo Wang, Xikun Zhang, James Zou, Katrina Kalantar

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.10641 2025-07-16 cs.SE cs.AI cs.CL 62%

A Code Comprehension Benchmark for Large Language Models for Code

Jayant Havare, Saurav Chaudhary, Ganesh Ramakrishnan, Kaushik Maharajan, Srikanth Tamilselvam

机构 * IIT Bombay(印度理工学院博伊斯分校) IBM Research(IBM研究院)

专题命中 安全评测 :DPO(abstract);分类 cs.CL、cs.AI

Comments 10 Pages, 5 Figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.10920 2025-07-16 cs.CL cs.AI 62%

HanjaBridge: Resolving Semantic Ambiguity in Korean LLMs via Hanja-Augmented Pre-Training

Seungho Choi

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.10300 2025-07-15 cs.CV cs.AI cs.CL 62%

FaceLLM: A Multimodal Large Language Model for Face Understanding

Hatef Otroshi Shahreza, Sébastien Marcel

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL、cs.AI

Comments Accepted in ICCV 2025 workshops

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.10454 2025-07-15 cs.LG cs.AI 62%

An Interoperable Machine Learning Pipeline for Pediatric Obesity Risk Estimation

Hamed Fayyaz, Mehak Gupta, Alejandra Perez Ramirez, Claudine Jurkovitz, H. Timothy Bunnell, Thao-Ly T. Phan, Rahmatollah Beheshti

机构 * University of Delaware(德克萨斯大学) Southern Methodist University(南方 Methodist 大学) Nemours Children’s Health(Nemours 儿童健康)

专题命中 安全评测 :alignment(abstract);分类 cs.AI、cs.LG

Comments This paper has been accepted in Machine Learning for Health (ML4H) Symposium. Link: https://proceedings.mlr.press/v259/fayyaz25a.html

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.07611 2025-07-15 cs.CL cs.AI 62%

Knowledge-Augmented Multimodal Clinical Rationale Generation for Disease Diagnosis with Small Language Models

Shuai Niu, Jing Ma, Hongzhan Lin, Liang Bai, Zhihua Wang, Yida Xu, Yunya Song, Xian Yang

机构 * Hong Kong Baptist University(香港 Baptist 大学) Shanxi University(山西大学) Shanghai Institute for Advanced Study of Zhejiang University(浙江大学上海先进研究院) Hong Kong University of Science and Technology(香港科技大学) The University of Manchester(曼彻斯特大学)

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL、cs.AI

Comments 13 pages. 7 figures

Journal ref This paper is accpeted by ACL2025(Main)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.08719 2025-07-14 cs.CL cs.AI cs.SE 62%

Multilingual Multimodal Software Developer for Code Generation

Linzheng Chai, Jian Yang, Shukai Liu, Wei Zhang, Liran Wang, Ke Jin, Tao Sun, Congnan Liu, Chenchen Zhang, Hualei Zhu, Jiaheng Liu, Xianjie Wu, Ge Zhang, Tianyu Liu, Zhoujun Li

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

Comments Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.06715 2025-07-10 cs.CL cs.AI cs.IR 62%

CLI-RAG: A Retrieval-Augmented Framework for Clinically Structured and Context Aware Text Generation with LLMs

Garapati Keerthana, Manik Gupta

机构 * Birla Institute of Technology and Science, Pilani, Hyderabad, India(比拉理工学院和科学学院,海得拉巴,印度)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

Comments 12 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.06282 2025-07-10 cs.CR cs.AI cs.CL 62%

The bitter lesson of misuse detection

Hadrien Mariaccia, Charbel-Raphaël Segerie, Diego Dorn

机构 * Centre pour la Sécurité de l'IA (CeSIA)(人工智能安全研究中心) École Polytechnique Fédérale de Lausanne (EPFL)(洛桑联邦理工学院)

专题命中 安全评测 :jailbreak(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.06137 2025-07-09 cs.CL cs.AI cs.CV 62%

NeoBabel: A Multilingual Open Tower for Visual Generation

Mohammad Mahdi Derakhshani, Dheeraj Varghese, Marzieh Fadaee, Cees G. M. Snoek

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

Comments 34 pages, 12 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.05598 2025-07-09 cs.CL cs.AI 62%

Self-Review Framework for Enhancing Instruction Following Capability of LLM

Sihyun Park

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.05321 2025-07-09 cs.CY cs.AI 62%

AGACCI : Affiliated Grading Agents for Criteria-Centric Interface in Educational Coding Contexts

Kwangsuk Park, Jiwoong Yang

机构 * AA LAB, MODULABS(AA实验室,MODULABS) Aiffel, MODULABS(Aiffel,MODULABS) Department of Statistics, Inha university(统计系,Inha大学)

专题命中 安全评测 :alignment(abstract);分类 cs.AI、cs.CY

Comments Accepted at ICML 2025 Workshop on Multi-Agent Systems in the Era of Foundation Models: Opportunities, Challenges and Futures (MAS)

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.04945 2025-07-09 cs.CL cs.LG eess.SP 62%

MEIT: Multimodal Electrocardiogram Instruction Tuning on Large Language Models for Report Generation

Zhongwei Wan, Che Liu, Xin Wang, Chaofan Tao, Hui Shen, Jing Xiong, Rossella Arcucci, Huaxiu Yao, Mi Zhang

机构 * The Ohio State University(俄亥俄州立大学) Imperial College London(伦敦帝国理工学院) The University of Hong Kong(香港大学) University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.LG

Comments ACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.03175 2025-07-08 cs.LG cs.AI 62%

Understanding Knowledge Transferability for Transfer Learning: A Survey

Haohua Wang, Jingge Wang, Zijie Zhao, Yang Tan, Yanru Wu, Hanbing Liu, Jingyun Yang, Enming Zhang, Xiangyu Chen, Zhengze Rong, Shanxin Guo, Yang Li

机构 * Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院) Shenzhen Institute of Advanced Technology Chinese Academy of Sciences(中国科学院深圳先进技术研究院)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG

Comments 35 pages, 15 figures, submitted to ACM Computing Surveys

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.04428 2025-07-08 cs.CL cs.AI 62%

MoralBench: Moral Evaluation of LLMs

Jianchao Ji, Yutong Chen, Mingyu Jin, Wujiang Xu, Wenyue Hua, Yongfeng Zhang

机构 * Rutgers The State University of New Jersey(新泽西州立大学拉特格斯分校) University of Chicago(芝加哥大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

Comments Accepted to ACM SIGKDD Explorations Volume 27 Issue 1

详情

展开后加载摘要…

URL PDF HTML 收藏