arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 9434 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 9434 篇

2504.15125 2025-08-19 cs.AI 57%

Contemplative Artificial Intelligence

Ruben Laukkonen, Fionn Inglis, Shamil Chandaria, Lars Sandved-Smith, Edmundo Lopez-Sola, Jakob Hohwy, Jonathan Gold, Adam Elwood

专题命中 安全评测 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11848 2025-08-19 quant-ph cs.ET cs.LG 57%

Adversarial Robustness in Distributed Quantum Machine Learning

Pouya Kananian, Hans-Arno Jacobsen

机构 * Department of Electrical and Computer Engineering, University of Toronto(电气与计算机工程系,多伦多大学)

专题命中 安全评测 :trustworthy(abstract);分类 cs.LG

Comments This is a preprint of a book chapter that is planned to be published in "Quantum Robustness in Artificial Intelligence" by Springer Nature

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.03603 2025-08-19 cs.LG 57%

MUC: Machine Unlearning for Contrastive Learning with Black-box Evaluation

Yihan Wang, Yiwei Lu, Guojun Zhang, Franziska Boenisch, Adam Dziedzic, Yaoliang Yu, Xiao-Shan Gao

机构 * University of Waterloo(滑铁卢大学) University of Ottawa(渥太华大学) Alibaba(阿里巴巴) CISPA Helmholtz Center for Information Security(信息安全赫尔姆霍兹中心) Vector Institute(向量研究所) Academy of Mathematics and Systems Science, Chinese Academy of Sciences(中国科学院数学与系统科学研究院) University of Chinese Academy of Sciences(中国科学院大学)

专题命中 安全评测 :alignment(abstract);分类 cs.LG

Comments Published in TMLR

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.13954 2025-08-19 cs.CL 57%

Measuring Social Biases in Masked Language Models by Proxy of Prediction Quality

Rahul Zalkikar, Kanchan Chandra

机构 * New York University(纽约大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

Journal ref Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2025, pages 1337--1361

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11514 2025-08-18 cs.LG 57%

DiCriTest: Testing Scenario Generation for Decision-Making Agents Considering Diversity and Criticality

Qitong Chu, Yufeng Yue, Danya Yao, Huaxin Pei

机构 * School of Automation, Beijing Institute of Technology(北京理工大学自动化学院) Department of Automation, Tsinghua University(清华大学自动化系)

专题命中 安全评测 :safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10869 2025-08-15 cs.CV cs.AI 57%

Medico 2025: Visual Question Answering for Gastrointestinal Imaging

Sushant Gautam, Vajira Thambawita, Michael Riegler, Pål Halvorsen, Steven Hicks

机构 * SimulaMet - Simula Metropolitan Center for Digital Engineering, Oslo, Norway(SimulaMet - Simula Metropolitan Center for Digital Engineering,挪威奥斯陆) Simula Research Laboratory, Oslo, Norway(Simula研究实验室,挪威奥斯陆) OsloMet - Oslo Metropolitan University, Oslo, Norway(OsloMet - 奥斯陆 Metropolitan 大学,挪威奥斯陆)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.13307 2025-08-15 cs.CV cs.AI 57%

Quantitative Comparison of Fine-Tuning Techniques for Pretrained Latent Diffusion Models in the Generation of Unseen SAR Images

Solène Debuysère, Nicolas Trouvé, Nathan Letheule, Olivier Lévêque, Elise Colin

机构 * Paris-Saclay University(巴黎-萨克雷大学) ONERA - The French Aerospace Lab(法国航空航天实验室)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10397 2025-08-15 cs.CV cs.AI 57%

PQ-DAF: Pose-driven Quality-controlled Data Augmentation for Data-scarce Driver Distraction Detection

Haibin Sun, Xinghui Song

机构 * College of Computer Science and Engineering, Shandong University of Science and Technology(计算机科学与工程学院,山东科技大学)

专题命中 安全评测 :safety(abstract);分类 cs.AI

Comments 11 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10358 2025-08-15 cs.AI 57%

What to Ask Next? Probing the Imaginative Reasoning of LLMs with TurtleSoup Puzzles

Mengtao Zhou, Sifan Wu, Huan Zhang, Qi Sima, Bang Liu

专题命中 安全评测 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.09893 2025-08-14 cs.AI 57%

RAGulating Compliance: A Multi-Agent Knowledge Graph for Regulatory QA

Bhavik Agarwal, Hemant Sunil Jomraj, Simone Kaplunov, Jack Krolick, Viktoria Rojkova

机构 * MasterControl AI Research(MasterControl AI研究院)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.17609 2025-08-14 cs.AI 57%

Integrating Visual Interpretation and Linguistic Reasoning for Math Problem Solving

Zixian Guo, Ming Liu, Qilong Wang, Zhilong Ji, Jinfeng Bai, Lei Zhang, Wangmeng Zuo

机构 * Harbin Institute of Technology(哈尔滨理工大学) The Hong Kong Polytechnic University(香港理工大学) Tianjin University(天津大学) Tomorrow Advancing Life(明天进步生命) Pazhou Lab, Guangzhou(广州琶洲实验室)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

Comments ICCV2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06555 2025-08-13 cs.CV cs.CY cs.MA 57%

StyleTailor: Towards Personalized Fashion Styling via Hierarchical Negative Feedback

Hongbo Ma, Fei Shen, Hongbin Xu, Xiaoce Wang, Gang Xu, Jinkai Zheng, Liangqiong Qu, Ming Li

专题命中 安全评测 :alignment(abstract);分类 cs.CY

Comments 24pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07163 2025-08-12 cs.RO cs.AI cs.NE 57%

Integrating Neurosymbolic AI in Advanced Air Mobility: A Comprehensive Survey

Kamal Acharya, Iman Sharifi, Mehul Lad, Liang Sun, Houbing Song

机构 * Department of Information Systems, University of Maryland, Baltimore County(信息系统系,马里兰大学巴尔的摩分校) Department of Mechanical and Aerospace Engineering, The George Washington University(机械与航空航天工程系,乔治华盛顿大学) Department of Mechanical Engineering, Baylor University(机械工程系,贝勒大学)

专题命中 安全评测 :safety(abstract);分类 cs.AI

Comments 9 pages, 4 figures, IJCAI-2025 (accepted)

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.17632 2025-08-12 cs.AI cs.CV cs.MM 57%

D-Judge: How Far Are We? Assessing the Discrepancies Between AI-synthesized and Natural Images through Multimodal Guidance

Renyang Liu, Ziyu Lyu, Wei Zhou, See-Kiong Ng

机构 * School of Cyber Science and Technology, Shenzhen Campus of Sun Yat-sen University(中山大学信息科学与技术学院(深圳校区)) Institute of Data Science, National University of Singapore(新加坡国立大学数据科学研究所) College of Modern Engineering and the Engineering Research Center of Cyberspace, Yunnan University(云南大学现代工程学院及空天信息工程研究中心)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

Comments Accepted by ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.11569 2025-08-12 eess.IV cs.AI cs.CV 57%

Are Vision Foundation Models Ready for Out-of-the-Box Medical Image Registration?

Hanxue Gu, Yaqian Chen, Nicholas Konz, Qihang Li, Maciej A. Mazurowski

机构 * Department of Electrical and Computer Engineering, Duke University(电子工程与计算机科学系,杜克大学) Department of Biostatistics and Bioinformatics, Duke University(生物统计学与生物信息学系,杜克大学) Departments of Biostatistics and Bioinformatics, Radiology, Electrical and Computer Engineering, and Computer Science, Duke University(生物统计学与生物信息学系、放射学、电子工程与计算机科学系,杜克大学)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

Comments 3 figures, 9 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.17590 2025-08-12 cs.CV cs.AI cs.RO 57%

DRAMA-X: A Fine-grained Intent Prediction and Risk Reasoning Benchmark For Driving

Mihir Godbole, Xiangbo Gao, Zhengzhong Tu

机构 * Texas A&M University(德克萨斯A&M大学)

专题命中 安全评测 :safety(abstract);分类 cs.AI

Comments 19 pages, 5 figures, Preprint under review. Code available at: https://github.com/taco-group/DRAMA-X

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06155 2025-08-11 cs.CL 57%

Semantic and Structural Analysis of Implicit Biases in Large Language Models: An Interpretable Approach

Renhan Zhang, Lian Lian, Zhen Qi, Guiran Liu

专题命中 安全评测 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05987 2025-08-11 cs.CL 57%

Adversarial Topic-aware Prompt-tuning for Cross-topic Automated Essay Scoring

Chunyun Zhang, Hongyan Zhao, Chaoran Cui, Qilong Song, Zhiqing Lu, Shuai Gong, Kailin Liu

机构 * Shandong University of Finance and Economics(山东财经大学) University of Toronto(多伦多大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.08947 2025-08-11 cs.CL 57%

Structured Convergence in Large Language Model Representations via Hierarchical Latent Space Folding

Fenella Harcourt, Naderdel Piero, Gilbert Sutherland, Daphne Holloway, Harriet Bracknell, Julian Ormsby

专题命中 安全评测 :alignment(abstract);分类 cs.CL

Comments arXiv admin note: This paper has been withdrawn by arXiv due to disputed and unverifiable authorship

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05687 2025-08-11 cs.MA cs.AI 57%

Risk Analysis Techniques for Governed LLM-based Multi-Agent Systems

Alistair Reid, Simon O'Callaghan, Liam Carroll, Tiberio Caetano

机构 * Gradient Institute Ltd.(梯度研究所有限公司)

专题命中 安全评测 :red teaming(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05508 2025-08-08 cs.AI 57%

Auto-Eval Judge: Towards a General Agentic Framework for Task Completion Evaluation

Roshita Bhonsle, Rishav Dutta, Sneha Vavilapalli, Harsh Seth, Abubakarr Jaye, Yapei Chang, Mukund Rungta, Emmanuel Aboah Boateng, Sadid Hasan, Ehi Nosakhare, Soundar Srinivasan

机构 * University of Massachusetts Amherst(马萨诸塞大学阿姆赫斯特分校) Microsoft Corporation(微软公司) University of Maryland(马里兰大学)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05468 2025-08-08 cs.CL 57%

TASE: Token Awareness and Structured Evaluation for Multilingual Language Models

Chenzhuo Zhao, Xinda Wang, Yue Huang, Junting Lu, Ziqian Liu

专题命中 安全评测 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05237 2025-08-08 cs.CV cs.AI 57%

Navigating the Trade-off: A Synthesis of Defensive Strategies for Zero-Shot Adversarial Robustness in Vision-Language Models

Zane Xu, Jason Sun

专题命中 安全评测 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05083 2025-08-08 cs.AI 57%

MedMKEB: A Comprehensive Knowledge Editing Benchmark for Medical Multimodal Large Language Models

Dexuan Xu, Jieyi Wang, Zhongyan Chai, Yongzhi Cao, Hanpin Wang, Huamin Zhang, Yu Huang

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

Comments 18 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04469 2025-08-07 cs.CV cs.CL 57%

FrEVL: Leveraging Frozen Pretrained Embeddings for Efficient Vision-Language Understanding

Emmanuelle Bourigault, Pauline Bourigault

机构 * Department of Engineering Science, University of Oxford(牛津大学工程科学系) Department of Electrical Engineering, Imperial College London(伦敦帝国学院电子工程系)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

Comments 8 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04199 2025-08-07 cs.CL 57%

Reasoning Beyond Labels: Measuring LLM Sentiment in Low-Resource, Culturally Nuanced Contexts

Millicent Ochieng, Anja Thieme, Ignatius Ezeani, Risa Ueno, Samuel Maina, Keshet Ronen, Javier Gonzalez, Jacki O'Neill

机构 * Microsoft Research(微软研究院) Lancaster University(兰卡斯特大学) University of Washington(华盛顿大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04011 2025-08-07 cs.HC cs.AI 57%

StepWrite: Adaptive Planning for Speech-Driven Text Generation

Hamza El Alaoui, Atieh Taheri, Yi-Hao Peng, Jeffrey P. Bigham

机构 * School of Computer Science Carnegie Mellon University Pittsburgh PA USA School of Computer Science Carnegie Mellon University

专题命中 安全评测 :alignment(abstract);分类 cs.AI

Comments This paper has been accepted to UIST 2025. For additional materials and project details, please see: https://www.cs.cmu.edu/~helalaou/publications/stepwrite

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03734 2025-08-07 eess.IV cs.AI cs.CV 57%

A Survey of Multimodal Ophthalmic Diagnostics: From Task-Specific Approaches to Foundational Models

Xiaoling Luo, Ruli Zheng, Qiaojian Zheng, Zibo Du, Shuo Yang, Meidan Ding, Qihao Xu, Chengliang Liu, Linlin Shen

机构 * College of Computer Science and Software Engineering, Shenzhen University, Shenzhen, China(深圳大学计算机科学与软件工程学院) Shenzhen Key Laboratory of Visual Object Detection and Recognition, Harbin Institute of Technology, Shenzhen, 518055, China(视觉对象检测与识别深圳重点实验室) Laboratory for Artificial Intelligence in Design, Hong Kong(人工智能设计实验室) School of Artificial Intelligence, Shenzhen University, Shenzhen, China(深圳大学人工智能学院)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03722 2025-08-07 cs.CV cs.AI 57%

Multimodal Video Emotion Recognition with Reliable Reasoning Priors

Zhepeng Wang, Yingjian Zhu, Guanghao Dong, Hongzhu Yi, Feng Chen, Xinming Wang, Jun Xie

机构 * Lenovo Research(联想研究院) School of Artificial Intelligence, UCAS(中国科学院大学人工智能学院) Institute of Automation, CAS(中国科学院自动化研究所) Macau University of Science and Technology(澳门科学理工学院) School of Computer Science and Technology, UCAS(中国科学院大学计算机科学与技术学院)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

Comments preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.13783 2025-08-07 cs.MA cs.AI cs.GT cs.SY eess.SY 57%

A Value Based Parallel Update MCTS Method for Multi-Agent Cooperative Decision Making of Connected and Automated Vehicles

Ye Han, Lijun Zhang, Dejian Meng, Zhuang Zhang, Xingyu Hu, Songyu Weng

机构 * School of Automotive Studies, Tongji University(同济大学汽车学院)

专题命中 安全评测 :safety(abstract);分类 cs.AI

Comments arXiv admin note: text overlap with arXiv:2408.04295 by other authors

详情

展开后加载摘要…

URL PDF HTML 收藏