arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-08-20 至 2025-08-20 共收录 8 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 8 篇

2508.13179 2025-08-20 cs.CY cs.AI 88%

Toward an African Agenda for AI Safety

Samuel T. Segun, Rachel Adams, Ana Florido, Scott Timcke, Jonathan Shock, Leah Junck, Fola Adeleke, Nicolas Grossman, Ayantola Alayande, Jerry John Kponyo, Matthew Smith, Dickson Marfo Fosu, Prince Dawson Tetteh, Juliet Arthur, Stephanie Kasaon, Odilile Ayodele, Laetitia Badolo, Paul Plantinga, Michael Gastrow, Sumaya Nur Adan, Joanna Wiaterek, Cecil Abungu, Kojo Apeagyei, Luise Eder, Tegawende Bissyande

专题命中 安全评测 :safety(title,abstract);AI safety(title,abstract);分类 cs.AI、cs.CY

Comments 28 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.13787 2025-08-20 cs.MA cs.AI cs.NI 79%

BetaWeb: Towards a Blockchain-enabled Trustworthy Agentic Web

Zihan Guo, Yuanjian Zhou, Chenyi Wang, Linlin You, Minjie Bian, Weinan Zhang

机构 * Shanghai Innovation Institute(上海创新研究院) Sun Yat-sen University(中山大学) Zhejiang University(浙江大学) Shanghai Data Group Co., Ltd(上海数据集团有限公司) Shanghai Jiao Tong University(上海交通大学)

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI

Comments A technical report with 21 pages, 3 figures, and 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.13180 2025-08-20 cs.AI cs.LG 62%

Search-Time Data Contamination

Ziwen Han, Meher Mankikar, Julian Michael, Zifan Wang

机构 * Scale AI

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.13465 2025-08-20 cs.AI 57%

LM Agents May Fail to Act on Their Own Risk Knowledge

Yuzhi Tang, Tianxiao Li, Elizabeth Li, Chris J. Maddison, Honghua Dong, Yangjun Ruan

专题命中 安全评测 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.03666 2025-08-20 cs.RO cs.LG 57%

Hybrid Machine Learning Model with a Constrained Action Space for Trajectory Prediction

Alexander Fertig, Lakshman Balasubramanian, Michael Botsch

机构 * Technische Hochschule Ingolstadt, AImotion Bavaria(图腾工业大学,拜耳巴伐利亚人工智能公司) Technische Hochschule Ingolstadt, Research Center CARISSMA(图腾工业大学,CARISSMA研究中心)

专题命中 安全评测 :safety(abstract);分类 cs.LG

Comments Copyright 2025 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works

Journal ref 2025 IEEE Intelligent Vehicles Symposium (IV)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.13880 2025-08-20 cs.CV 50%

In-hoc Concept Representations to Regularise Deep Learning in Medical Imaging

Valentina Corbetta, Floris Six Dijkstra, Regina Beets-Tan, Hoel Kervadec, Kristoffer Wickstrøm, Wilson Silva

机构 * The Netherlands Cancer Institute(荷兰癌症研究所) Utrecht University(乌得勒支大学) Maastricht University(马斯特里赫特大学) University of Amsterdam(阿姆斯特丹大学) Amsterdam UMC(阿姆斯特丹大学医学中心) UiT The Arctic University of Norway(挪威北莫斯堡大学)

专题命中 安全评测 :trustworthy(abstract)

Comments 13 pages, 13 figures, 2 tables, accepted at PHAROS-AFE-AIMI Workshop in conjunction with the International Conference on Computer Vision (ICCV), 2025. This is the submitted manuscript with added link to the github repo, funding acknowledgments and author names and affiliations, and a correction to numbers in Table 1. Final version not published yet

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.13413 2025-08-20 cs.HC cs.SE 50%

Large Language Models as Visualization Agents for Immersive Binary Reverse Engineering

Dennis Brown, Samuel Mulder

专题命中 安全评测 :alignment(abstract)

Comments Accepted to IEEE VISSOFT 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06189 2025-08-20 cs.CV 50%

MA-CBP: A Criminal Behavior Prediction Framework Based on Multi-Agent Asynchronous Collaboration

Cheng Liu, Daou Zhang, Tingxu Liu, Yuhan Wang, Jinyang Chen, Yuexuan Li, Xinying Xiao, Chenbo Xin, Ziru Wang, Weichao Wu

专题命中 安全评测 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏