arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-08-28 至 2025-08-28 共收录 11 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 11 篇

2508.12733 2025-08-28 cs.CL cs.AI 84%

LinguaSafe: A Comprehensive Multilingual Safety Benchmark for Large Language Models

Zhiyuan Ning, Tianle Gu, Jiaxin Song, Shixin Hong, Lingyu Li, Huacan Liu, Jie Li, Yixu Wang, Meng Lingyu, Yan Teng, Yingchun Wang

机构 * Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)

专题命中 安全评测 :safety(title,abstract);alignment(abstract);分类 cs.CL、cs.AI

Comments 7pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.02531 2025-08-28 cs.CY cs.CL 81%

Towards New Benchmark for AI Alignment & Sentiment Analysis in Socially Important Issues: A Comparative Study of Human and LLMs in the Context of AGI

Ljubisa Bojic, Dylan Seychell, Milan Cabarkapa

专题命中 安全评测 :alignment(title,abstract);分类 cs.CL、cs.CY

Comments 34 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.19487 2025-08-28 cs.LG cs.AI 62%

Data-Efficient Symbolic Regression via Foundation Model Distillation

Wangyang Ying, Jinghan Zhang, Haoyue Bai, Nanxu Gong, Xinyuan Wang, Kunpeng Liu, Chandan K. Reddy, Yanjie Fu

机构 * Institute for Clarity in Documentation(文档清晰研究所) Inria Paris-Rocquencourt(巴黎-罗克琴特研究所) Rajiv Gandhi University(拉吉夫·甘地大学) Tsinghua University(清华大学) Palmer Research Laboratories(帕尔默研究实验室) Arizona State University(亚利桑那州立大学) Clemson University(克莱姆森大学)

专题命中 安全评测 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.19271 2025-08-28 cs.CL cs.AI 62%

Rethinking Reasoning in LLMs: Neuro-Symbolic Local RetoMaton Beyond ICL and CoT

Rushitha Santhoshi Mamidala, Anshuman Chhabra, Ankur Mali

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.19980 2025-08-28 cs.LG 57%

Evaluating Language Model Reasoning about Confidential Information

Dylan Sam, Alexander Robey, Andy Zou, Matt Fredrikson, J. Zico Kolter

专题命中 安全评测 :safety(abstract);分类 cs.LG

Comments 20 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.19882 2025-08-28 cs.SE cs.AI 57%

Generative AI for Testing of Autonomous Driving Systems: A Survey

Qunying Song, He Ye, Mark Harman, Federica Sarro

机构 * University College London(伦敦大学学院)

专题命中 安全评测 :safety(abstract);分类 cs.AI

Comments 67 pages, 6 figures, 29 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.19641 2025-08-28 cs.CR cs.AI 57%

Intellectual Property in Graph-Based Machine Learning as a Service: Attacks and Defenses

Lincan Li, Bolin Shen, Chenxi Zhao, Yuxiang Sun, Kaixiang Zhao, Shirui Pan, Yushun Dong

机构 * Department of Computer Science, Florida State University(佛罗里达州立大学计算机科学系) Northeastern University(东北大学) University of Wisconsin-Madison(威斯康星大学麦迪逊分校) University of Notre Dame(诺丁汉大学) School of Information and Communication Technology, Griffith University(格里菲斯大学信息与通信技术学院)

专题命中 安全评测 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.09242 2025-08-28 cs.AI 57%

From large language models to multimodal AI: A scoping review on the potential of generative AI in medicine

Lukas Buess, Matthias Keicher, Nassir Navab, Andreas Maier, Soroosh Tayebi Arasteh

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

Journal ref Biomed. Eng. Lett. 15 (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.19773 2025-08-28 cs.CV 50%

The Return of Structural Handwritten Mathematical Expression Recognition

Jakob Seitz, Tobias Lengfeld, Radu Timofte

机构 * Computer Vision Lab, CAIDAS \& IFI, University of W\"urzburg, W\"urzburg, Germany

专题命中 安全评测 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.19639 2025-08-28 cs.MM 50%

FakeSV-VLM: Taming VLM for Detecting Fake Short-Video News via Progressive Mixture-Of-Experts Adapter

Junxi Wang, Yaxiong Wang, Lechao Cheng, Zhun Zhong

专题命中 安全评测 :alignment(abstract)

Comments EMNLP2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21696 2025-08-28 eess.SP 50%

Edge Agentic AI Framework for Autonomous Network Optimisation in O-RAN

Abdelaziz Salama, Zeinab Nezami, Mohammed M. H. Qazzaz, Maryam Hafeez, Syed Ali Raza Zaidi

专题命中 安全评测 :safety(abstract)

Journal ref IEEE International Symposium on Personal, Indoor and Mobile Radio Communications (PIMRC), 2025

详情

展开后加载摘要…

URL PDF HTML 收藏