arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-09-24 至 2025-09-24 共收录 10 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 10 篇

2501.01346 2025-09-24 cs.CV cs.CL 79%

Large Vision-Language Model Alignment and Misalignment: A Survey Through the Lens of Explainability

Dong Shu, Haiyan Zhao, Jingyu Hu, Weiru Liu, Ali Payani, Lu Cheng, Mengnan Du

机构 * Northwestern University(西北大学) New Jersey Institute of Technology(新泽西理工学院) University of Bristol(布里斯托大学) Cisco Research(思科研究) University of Illinois Chicago(伊利诺伊大学芝加哥分校)

专题命中 安全评测 :alignment(title,abstract);分类 cs.CL

Comments EMNLP 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15260 2025-09-24 cs.CL 79%

Toxicity Red-Teaming: Benchmarking LLM Safety in Singapore's Low-Resource Languages

Yujia Hu, Ming Shan Hee, Preslav Nakov, Roy Ka-Wei Lee

机构 * Singapore University of Technology and Design(新加坡科技设计大学) Mohamed bin Zayed University of Artificial Intelligence(马尔代夫 bin Zayed 人工智能大学)

专题命中 安全评测 :safety(title,abstract);分类 cs.CL

Comments 9 pages, EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19120 2025-09-24 cs.LG cs.AI cs.DC 76%

FedFiTS: Fitness-Selected, Slotted Client Scheduling for Trustworthy Federated Learning in Healthcare AI

Ferdinand Kahenga, Antoine Bagula, Sajal K. Das, Patrick Sello

机构 * Department of Computer Science University of the Western Cape(计算机科学系,西开普敦大学) Department of Computer Science Missouri University of Science and Technology(计算机科学系,密苏里科学与技术大学)

专题命中 安全评测 :trustworthy(title);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18557 2025-09-24 cs.AI 70%

LLMZ+: Contextual Prompt Whitelist Principles for Agentic LLMs

Tom Pawelek, Raj Patel, Charlotte Crowell, Noorbakhsh Amiri, Sudip Mittal, Shahram Rahimi, Andy Perkins

机构 * Department of Computer Science Mississippi State University(计算机科学系密苏里州立大学) Department of Computer Science The University of Alabama(计算机科学系阿拉巴马大学) Mississippi State University, Mississippi State, MS, USA(密苏里州立大学) The University of Alabama, Tuscaloosa, AL, USA(阿拉巴马大学)

专题命中 安全评测 :jailbreak(abstract);prompt injection(abstract);分类 cs.AI

Comments 7 pages, 5 figures, to be published and presented at ICMLA 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14944 2025-09-24 cs.HC 67%

LEKIA: Expert-Aligned AI Behavior Design for High-Risk Human-AI Interactions

Boning Zhao, Yutong Hu, Xinnuo Li

专题命中 安全评测 :alignment(abstract);safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18221 2025-09-24 cs.AI cs.LG 62%

Multimodal Health Risk Prediction System for Chronic Diseases via Vision-Language Fusion and Large Language Models

Dingxin Lu, Shurui Wu, Xinyi Huang

专题命中 安全评测 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.00046 2025-09-24 cs.CV cs.AI 57%

Leveraging Large Models to Evaluate Novel Content: A Case Study on Advertisement Creativity

Zhaoyi Joey Hou, Adriana Kovashka, Xiang Lorraine Li

机构 * Department of Computer Science University of Pittsburgh(计算机科学系宾夕法尼亚大学)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

Comments To Appear in EMNLP2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18568 2025-09-24 cs.LG 57%

Explainable Graph Neural Networks: Understanding Brain Connectivity and Biomarkers in Dementia

Niharika Tewari, Nguyen Linh Dan Le, Mujie Liu, Jing Ren, Ziqi Xu, Tabinda Sarwar, Veeky Baths, Feng Xia

机构 * School of Computing Technologies RMIT University Melbourne VIC Australia Department of Biological Sciences Department of Computer Science \& Information Systems Birla Institute of Technology Institute of Innovation, Science Sustainability Federation University Australia Ballarat VIC Australia RMIT University Birla Institute of Technology Federation University Australia

专题命中 安全评测 :trustworthy(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18869 2025-09-24 cs.DC 50%

On The Reproducibility Limitations of RAG Systems

Baiqiang Wang, Dongfang Zhao, Nathan R Tallent, Luanzheng Guo

专题命中 安全评测 :trustworthy(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17537 2025-09-24 cs.CV 50%

SimToken: A Simple Baseline for Referring Audio-Visual Segmentation

Dian Jin, Yanghao Zhou, Jinxing Zhou, Jiaqi Ma, Ruohao Guo, Dan Guo

专题命中 安全评测 :alignment(abstract)

Comments Project page: https://github.com/DianJin-HFUT/SimToken

详情

展开后加载摘要…

URL PDF HTML 收藏