arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 9434 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 9434 篇

2509.07039 2025-09-10 cs.LG cs.CV 57%

Benchmarking Vision Transformers and CNNs for Thermal Photovoltaic Fault Detection with Explainable AI Validation

Serra Aksoy

专题命中 安全评测 :alignment(abstract);分类 cs.LG

Comments 28 Pages, 4 Figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.07026 2025-09-10 cs.LO cs.AI 57%

Contradictions

Yang Xu, Shuwei Chen, Xiaomei Zhong, Jun Liu, Xingxing He

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

Comments 37 Pages,9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.07010 2025-09-10 cs.CV cs.AI cs.ET 57%

Human-in-the-Loop: Quantitative Evaluation of 3D Models Generation by Large Language Models

Ahmed R. Sadik, Mariusz Bujny

机构 * Honda Research Institute Europe - Germany(本田欧洲研究机构)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.06409 2025-09-09 cs.AI 57%

Teaching AI Stepwise Diagnostic Reasoning with Report-Guided Chain-of-Thought Learning

Yihong Luo, Wenwu He, Zhuo-Xu Cui, Dong Liang

机构 * Fujian University of Technology(福建工程学院) Fujian Provincial Key Laboratory of Big Data Mining and Applications(福建省大数据挖掘与应用重点实验室) Key Laboratory of Biomedical Imaging Science and System, Chinese Academy of Sciences(生物医学成像科学与系统重点实验室,中国科学院)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.06330 2025-09-09 cs.LG 57%

Exploring approaches to computational representation and classification of user-generated meal logs

Guanlan Hu, Adit Anand, Pooja M. Desai, Iñigo Urteaga, Lena Mamykina

专题命中 安全评测 :alignment(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.06269 2025-09-09 cs.AI 57%

REMI: A Novel Causal Schema Memory Architecture for Personalized Lifestyle Recommendation Agents

Vishal Raman, Vijai Aravindh R, Abhijith Ragav

机构 * Radian Group Inc.(Radian集团) Sri Sivasubramaniya Nadar College Of Engineering(Sri Sivasubramaniya纳达尔工程学院) Amazon(亚马逊)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

Comments 8 pages, 2 figures, Accepted at the OARS Workshop, KDD 2025, Paper link: https://oars-workshop.github.io/papers/Raman2025.pdf

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.05881 2025-09-09 cs.SE cs.AI 57%

GeoAnalystBench: A GeoAI benchmark for assessing large language models for spatial analysis workflow and code generation

Qianheng Zhang, Song Gao, Chen Wei, Yibo Zhao, Ying Nie, Ziru Chen, Shijie Chen, Yu Su, Huan Sun

机构 * University of Wisconsin-Madison(威斯康星大学麦迪逊分校) The Ohio State University(俄亥根州立大学)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

Comments 34 pages, 8 figures

Journal ref Transactions in GIS, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.05303 2025-09-09 cs.DC cs.AI 57%

Multi-IaC-Eval: Benchmarking Cloud Infrastructure as Code Across Multiple Formats

Sam Davidson, Li Sun, Bhavana Bhasker, Laurent Callot, Anoop Deoras

机构 * Amazon Web Services(亚马逊网络服务)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.06056 2025-09-09 cs.CL 57%

Synth-SBDH: A Synthetic Dataset of Social and Behavioral Determinants of Health for Clinical Text

Avijit Mitra, Zhichao Yang, Emily Druhl, Raelene Goodwin, Hong Yu

机构 * Manning College of Information and Computer Sciences, University of Massachusetts Amherst(马萨诸塞大学阿姆赫斯特分校信息与计算机科学学院) U.S. Department of Veterans Affairs(美国退伍军人事务部) Department of Medicine, University of Massachusetts Chan Medical School(马萨诸塞大学查恩医学院医学部) Miner School of Computer and Information Sciences, University of Massachusetts Lowell(马萨诸塞大学洛厄尔分校米纳尔计算机与信息科学学院)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

Comments Accepted at EMNLP 2025 (main) Github: https://github.com/avipartho/Synth-SBDH

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.12928 2025-09-09 cs.DL cs.AI cs.CV 57%

A Literature Review of Literature Reviews in Pattern Analysis and Machine Intelligence

Penghai Zhao, Xin Zhang, Jiayue Cao, Ming-Ming Cheng, Jian Yang, Xiang Li

机构 * VCIP, CS, Nankai University, Tianjin 300000 , China(VCIP、计算机科学系、南开大学、天津)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

Comments V2, V3, and V4 with incremental quality improvements. V5, V6 introduce major updates

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04794 2025-09-08 cs.CL 57%

Personality as a Probe for LLM Evaluation: Method Trade-offs and Downstream Effects

Gunmay Handa, Zekun Wu, Adriano Koshiyama, Philip Treleaven

机构 * University College London(伦敦大学学院)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04292 2025-09-05 cs.CL 57%

Inverse IFEval: Can LLMs Unlearn Stubborn Training Conventions to Follow Real Instructions?

Qinyan Zhang, Xinping Lei, Ruijie Miao, Yu Fu, Haojie Fan, Le Chang, Jiafan Hou, Dingling Zhang, Zhongfei Hou, Ziqiang Yang, Changxin Pu, Fei Hu, Jingkai Liu, Mengyun Liu, Yang Liu, Xiang Gao, Jiaheng Liu, Tong Yang, Zaiyuan Wang, Ge Zhang, Wenhao Huang

机构 * ByteDance Seed(字节跳动种子) Nanjing University(南京大学) Peking University(北京大学) Beijing University of Posts and Telecommunications(北京邮电大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04053 2025-09-05 cs.LG 57%

On Aligning Prediction Models with Clinical Experiential Learning: A Prostate Cancer Case Study

Jacqueline J. Vallon, William Overman, Wanqiao Xu, Neil Panjwani, Xi Ling, Sushmita Vij, Hilary P. Bagshaw, John T. Leppert, Sumit Shah, Geoffrey Sonn, Sandy Srinivas, Erqi Pollom, Mark K. Buyyounouski, Mohsen Bayati

专题命中 安全评测 :alignment(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.03787 2025-09-05 cs.IR cs.CL 57%

Evaluating the Robustness of Retrieval-Augmented Generation to Adversarial Evidence in the Health Domain

Shakiba Amirshahi, Amin Bigdeli, Charles L. A. Clarke, Amira Ghenai

机构 * University of Waterloo(滑铁卢大学) Toronto Metropolitan University(多伦多 Metropolitan 大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.03565 2025-09-05 cs.CL cs.MM 57%

ResearchPulse: Building Method-Experiment Chains through Multi-Document Scientific Inference

Qi Chen, Jingxuan Wei, Zhuoya Yao, Haiguang Wang, Gaowei Wu, Bihui Yu, Siyuan Li, Cheng Tan

机构 * University of Chinese Academy of Sciences(中国科学院大学) Zhejiang University(浙江大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

Comments Accepted to ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.03383 2025-09-04 cs.AI cs.RO 57%

ANNIE: Be Careful of Your Robots

Yiyang Huang, Zixuan Wang, Zishen Wan, Yapeng Tian, Haobo Xu, Yinhe Han, Yiming Gan

机构 * Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) Georgia Institute of Technology(佐治亚理工学院) University of Texas at Dallas(德克萨斯大学达拉斯分校)

专题命中 安全评测 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.02820 2025-09-04 cs.LG 57%

Unlearning That Lasts: Utility-Preserving, Robust, and Almost Irreversible Forgetting in LLMs

Naman Deep Singh, Maximilian Müller, Francesco Croce, Matthias Hein

机构 * University of Tübingen & Tübingen AI Center, Germany(图宾根大学及图宾根人工智能中心) EPFL, Switzerland(瑞士联邦理工学院)

专题命中 安全评测 :safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.18190 2025-09-03 cs.AI cs.DB cs.IR 57%

ST-Raptor: LLM-Powered Semi-Structured Table Question Answering

Zirui Tang, Boyu Niu, Xuanhe Zhou, Boxiu Li, Wei Zhou, Jiannan Wang, Guoliang Li, Xinyi Zhang, Fan Wu

机构 * Shanghai Jiao Tong University(上海交通大学) Simon Fraser University(西蒙弗雷泽大学) Tsinghua University(清华大学) Renmin University of China(中国人民大学)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

Comments Extension of our SIGMOD 2026 paper. Please refer to source code available at: https://github.com/weAIDB/ST-Raptor

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.09889 2025-09-03 cs.AI 57%

Profile-Aware Maneuvering: A Dynamic Multi-Agent System for Robust GAIA Problem Solving by AWorld

Zhitian Xie, Qintong Wu, Chengyue Yu, Chenyi Zhuang, Jinjie Gu

机构 * AWorld Team, Inclusion AI(Inclusion AI团队)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.20435 2025-09-03 cs.CY 57%

Assessing the Sustainability and Trustworthiness of Federated Learning Models

Chao Feng, Alberto Huertas Celdran, Pedro Miguel Sanchez Sanchez, Lynn Zumtaugwald, Gerome Bovet, Burkhard Stiller

专题命中 安全评测 :trustworthy(abstract);分类 cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00936 2025-09-03 cs.AI 57%

UrbanInsight: A Distributed Edge Computing Framework with LLM-Powered Data Filtering for Smart City Digital Twins

Kishor Datta Gupta, Md Manjurul Ahsan, Mohd Ariful Haque, Roy George, Azmine Toushik Wasi

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00425 2025-09-03 cs.CL 57%

The Gold Medals in an Empty Room: Diagnosing Metalinguistic Reasoning in LLMs with Camlang

Fenghua Liu, Yulong Chen, Yixuan Liu, Zhujun Jin, Solomon Tsai, Ming Zhong

机构 * University of Cambridge(剑桥大学) University of Oxford(牛津大学) UIUC(伊利诺伊大学香槟分校)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

Comments Working in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00029 2025-09-03 cs.SD cs.AI cs.MM eess.AS 57%

From Sound to Sight: Towards AI-authored Music Videos

Leo Vitasovic, Stella Graßhof, Agnes Mercedes Kloft, Ville V. Lehtola, Martin Cunneen, Justyna Starostka, Glenn McGarry, Kun Li, Sami S. Brandt

机构 * IT University of Copenhagen(哥本哈根IT大学) Aalto University(艾尔沃斯大学) University of Twente(特文特大学) University of Limerick(利默里克大学) University of Nottingham(诺丁汉大学)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

Comments 1st Workshop on Generative AI for Storytelling (AISTORY), 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00010 2025-09-03 physics.ao-ph cs.LG 57%

CERA: A Framework for Improved Generalization of Machine Learning Models to Changed Climates

Shuchang Liu, Paul A. O'Gorman

专题命中 安全评测 :alignment(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.06794 2025-09-03 cs.CV cs.CL 57%

Does Acceleration Cause Hidden Instability in Vision Language Models? Uncovering Instance-Level Divergence Through a Large-Scale Empirical Study

Yizheng Sun, Hao Li, Chang Xu, Hongpeng Zhou, Chenghua Lin, Riza Batista-Navarro, Jingyuan Sun

机构 * University of Manchester(曼彻斯特大学) Microsoft Research(微软研究院)

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL

Comments Accepted to EMNLP 2025 Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20288 2025-08-29 eess.SY cs.LG cs.SY 57%

Neural Spline Operators for Risk Quantification in Stochastic Systems

Zhuoyuan Wang, Raffaele Romagnoli, Kamyar Azizzadenesheli, Yorie Nakahira

机构 * Department of Electrical and Computering Engineering, Carnegie Mellon University(电气与计算机工程系,卡内基梅隆大学) School of Science and Engineering, Department of Mathematics and Computer Science, Duquesne University(科学与工程学院,数学与计算机科学系,杜克森大学)

专题命中 安全评测 :safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.19980 2025-08-28 cs.LG 57%

Evaluating Language Model Reasoning about Confidential Information

Dylan Sam, Alexander Robey, Andy Zou, Matt Fredrikson, J. Zico Kolter

专题命中 安全评测 :safety(abstract);分类 cs.LG

Comments 20 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.19882 2025-08-28 cs.SE cs.AI 57%

Generative AI for Testing of Autonomous Driving Systems: A Survey

Qunying Song, He Ye, Mark Harman, Federica Sarro

机构 * University College London(伦敦大学学院)

专题命中 安全评测 :safety(abstract);分类 cs.AI

Comments 67 pages, 6 figures, 29 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.19641 2025-08-28 cs.CR cs.AI 57%

Intellectual Property in Graph-Based Machine Learning as a Service: Attacks and Defenses

Lincan Li, Bolin Shen, Chenxi Zhao, Yuxiang Sun, Kaixiang Zhao, Shirui Pan, Yushun Dong

机构 * Department of Computer Science, Florida State University(佛罗里达州立大学计算机科学系) Northeastern University(东北大学) University of Wisconsin-Madison(威斯康星大学麦迪逊分校) University of Notre Dame(诺丁汉大学) School of Information and Communication Technology, Griffith University(格里菲斯大学信息与通信技术学院)

专题命中 安全评测 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.09242 2025-08-28 cs.AI 57%

From large language models to multimodal AI: A scoping review on the potential of generative AI in medicine

Lukas Buess, Matthias Keicher, Nassir Navab, Andreas Maier, Soroosh Tayebi Arasteh

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

Journal ref Biomed. Eng. Lett. 15 (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏