arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-08-01 至 2025-08-01 共收录 9 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 9 篇

2507.10817 2025-08-01 stat.AP 78%

Is Your Model Risk ALARP? Evaluating Prospective Safety-Critical Applications of Complex Models

Domenic Di Francesco, Alan Forrest, Fiona McGarry, Nicholas Hall, Adam Sobey

专题命中 安全评测 :safety(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.23535 2025-08-01 cs.LG cs.AI cs.CY 67%

Transparent AI: The Case for Interpretability and Explainability

Dhanesh Ramachandram, Himanshu Joshi, Judy Zhu, Dhari Gandhi, Lucas Hartman, Ananya Raval

机构 * Vector Institute for Artificial Intelligence(向量人工智能研究所)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.CY、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.22958 2025-08-01 cs.CV cs.AI cs.LG 62%

CHECK-MAT: Checking Hand-Written Mathematical Answers for the Russian Unified State Exam

Ruslan Khrulev

机构 * Moscow State University(莫斯科国立大学)

专题命中 安全评测 :alignment(abstract);分类 cs.AI、cs.LG

Comments 15 pages, 3 figures, 10 tables. Code is available at: https://github.com/Karifannaa/Auto-check-EGE-math

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.07695 2025-08-01 cs.CL cs.AI 62%

KeyKnowledgeRAG (K^2RAG): An Enhanced RAG method for improved LLM question-answering capabilities

Hruday Markondapatnaikuni, Basem Suleiman, Abdelkarim Erradi, Shijing Chen

机构 * University of Sydney(悉尼大学) University of New South Wales(新南威尔士大学) Qatar University(卡塔尔大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

Comments 21 pages, 14 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.23000 2025-08-01 cs.LG cs.CV 57%

Planning for Cooler Cities: A Multimodal AI Framework for Predicting and Mitigating Urban Heat Stress through Urban Landscape Transformation

Shengao Yi, Xiaojiang Li, Wei Tu, Tianhong Zhao

机构 * Department of City and Regional Planning, University of Pennsylvania(城市与区域规划系,宾夕法尼亚大学) Guangdong Key Laboratory for Urban Informatics, Guangdong-Hong Kong-Macao Joint Laboratory for Smart Cities(广东城市信息关键实验室、粤港澳大湾区智慧城市联合实验室) Shenzhen Key Laboratory of Spatial Information Smart Sensing(深圳空间信息智能感知关键实验室) Department of Urban Informatics, School of Architecture and Urban Planning, Shenzhen University(城市信息系,深圳大学) College of Big Data and Internet, Shenzhen Technology University(大数据与互联网学院,深圳科技大学)

专题命中 安全评测 :alignment(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.17186 2025-08-01 cs.CL 57%

FinGAIA: A Chinese Benchmark for AI Agents in Real-World Financial Domain

Lingfeng Zeng, Fangqi Lou, Zixuan Wang, Jiajie Xu, Jinyi Niu, Mengping Li, Yifan Dong, Qi Qi, Wei Zhang, Ziwei Yang, Jun Han, Ruilun Feng, Ruiqi Hu, Lejie Zhang, Zhengbo Feng, Yicheng Ren, Xin Guo, Zhaowei Liu, Dongpo Cheng, Weige Cai, Liwen Zhang

专题命中 安全评测 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.12284 2025-08-01 cs.LG stat.ML 57%

GrokAlign: Geometric Characterisation and Acceleration of Grokking

Thomas Walker, Ahmed Imtiaz Humayun, Randall Balestriero, Richard Baraniuk

机构 * Department of Electrical and Computer Engineering, Rice University(电气与计算机工程系, Rice大学) Google Research(谷歌研究院) Department of Computer Science, Brown University(计算机科学系, Brown大学)

专题命中 安全评测 :alignment(abstract);分类 cs.LG

Comments 23 pages, 11 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.02123 2025-08-01 cs.CV cs.LG 57%

FovEx: Human-Inspired Explanations for Vision Transformers and Convolutional Neural Networks

Mahadev Prasad Panda, Matteo Tiezzi, Martina Vilas, Gemma Roig, Bjoern M. Eskofier, Dario Zanca

机构 * Department AIBE, FAU Erlangen-Nürnberg(FAU埃朗根-纽伦堡大学AIBE部门) PAVIS, Istituto Italiano di Tecnologia (IIT), Genova, Italy(意大利技术研究院(IIT)帕维斯部门,热那亚,意大利) Ernst Strüngmann Institute for Neuroscience, Frankfurt, Germany(神经科学埃朗根-纽伦堡研究所,法兰克福,德国) Goethe-Universität Frankfurt am Main(法兰克福大学) Institute of AI for Health, Helmholtz Zentrum München, Munich, Germany(健康人工智能研究所,海德堡中心慕尼黑,德国)

专题命中 安全评测 :alignment(abstract);分类 cs.LG

Comments Accepted in the International Journal of Computer Vision (Springer Nature)

Journal ref Int J Comput Vis (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.23489 2025-08-01 eess.SY cs.SY 50%

Distributionally Robust Cascading Risk Quantification in Multi-Agent Rendezvous: Effects of Time Delay and Network Connectivity

Vivek Pandey, Nader Motee

专题命中 安全评测 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏