arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-11-03 至 2025-11-03 共收录 9 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 9 篇

2510.27077 2025-11-03 cs.CL 85%

Contrastive Knowledge Transfer and Robust Optimization for Secure Alignment of Large Language Models

Jiasen Zheng, Huajun Zhang, Xu Yan, Ran Hao, Chong Peng

专题命中 安全评测 :alignment(title,abstract);safety(abstract);trustworthy(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20621 2025-11-03 cs.AI 79%

Towards the Formalization of a Trustworthy AI for Mining Interpretable Models explOiting Sophisticated Algorithms

Riccardo Guidotti, Martina Cinquini, Marta Marchiori Manerba, Mattia Setzu, Francesco Spinnato

机构 * University of Pisa(比萨大学)

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.12667 2025-11-03 cs.AI cs.LO 79%

Building Trustworthy AI by Addressing its 16+2 Desiderata with Goal-Directed Commonsense Reasoning

Alexis R. Tudor, Yankai Zeng, Huaduo Wang, Joaquin Arias, Gopal Gupta

机构 * University of Texas at Dallas(德克萨斯大学达拉斯分校)

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.27207 2025-11-03 cs.LG cs.AI 62%

Feature-Function Curvature Analysis: A Geometric Framework for Explaining Differentiable Models

Hamed Najafi, Dongsheng Luo, Jason Liu

机构 * Florida International University(佛罗里达国际大学)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.27521 2025-11-03 cs.HC cs.CY 57%

Independent Clinical Evaluation of General-Purpose LLM Responses to Signals of Suicide Risk

Nick Judd, Alexandre Vaz, Kevin Paeth, Layla Inés Davis, Milena Esherick, Jason Brand, Inês Amaro, Tony Rousmaniere

专题命中 安全评测 :alignment(abstract);分类 cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.27244 2025-11-03 cs.SE cs.AI 57%

Vintage Code, Modern Judges: Meta-Validation in Low Data Regimes

Ora Nova Fandina, Gal Amram, Eitan Farchi, Shmulik Froimovich, Raviv Gal, Wesam Ibraheem, Rami Katan, Alice Podolsky, Orna Raz

专题命中 安全评测 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.27065 2025-11-03 cs.LG cs.PF 57%

MLPerf Automotive

Radoyeh Shojaei, Predrag Djurdjevic, Mostafa El-Khamy, James Goel, Kasper Mecklenburg, John Owens, Pınar Muyan-Özçelik, Tom St. John, Jinho Suh, Arjun Suresh

机构 * University of California, Davis(加州大学戴维斯分校) Arm(ARM公司) Samsung(三星) Qualcomm(高通) California State University, Sacramento(加州州立大学萨克拉门托分校) Gilmet Labs(Gilmet实验室) NVIDIA(英伟达) AMD(超微半导体)

专题命中 安全评测 :safety(abstract);分类 cs.LG

Comments 16 pages, 5 figures, 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.26830 2025-11-03 cs.LG cs.CR 57%

SmoothGuard: Defending Multimodal Large Language Models with Noise Perturbation and Clustering Aggregation

Guangzhi Su, Shuchang Huang, Yutong Ke, Zhuohang Liu, Long Qian, Kaizhu Huang

机构 * Independent Researcher(独立研究者)

专题命中 安全评测 :safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.25761 2025-11-03 cs.CL 57%

DiagramEval: Evaluating LLM-Generated Diagrams via Graphs

Chumeng Liang, Jiaxuan You

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

Comments EMNLP 2025 Main

详情

展开后加载摘要…

URL PDF HTML 收藏