arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-08-14 至 2025-08-14 共收录 7 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 其他安全 7 篇

2508.09203 2025-08-14 cs.LG 79%

Building Safer Sites: A Large-Scale Multi-Level Dataset for Construction Safety Research

Zhenhui Ou, Dawei Li, Zhen Tan, Wenlin Li, Huan Liu, Siyuan Song

机构 * Arizona State University(亚利桑那州立大学)

专题命中 其他安全 :safety(title,abstract);分类 cs.LG

Comments The paper was accepted on the CIKM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.09963 2025-08-14 eess.SY cs.MA cs.RO cs.SY 78%

Online Safety under Multiple Constraints and Input Bounds using gatekeeper: Theory and Applications

Devansh R. Agrawal, Dimitra Panagou

机构 * Robotics Department, University of Michigan(密歇根大学机器人系)

专题命中 其他安全 :safety(title,abstract)

Comments 6 pages, 2 figures. Accepted for publication in IEEE L-CSS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.03012 2025-08-14 cs.AI cs.CL cs.CV 62%

Analyzing Finetuning Representation Shift for Multimodal LLMs Steering

Pegah Khayatan, Mustafa Shukor, Jayneel Parekh, Arnaud Dapogny, Matthieu Cord

机构 * ISIR, Sorbonne Université(ISIR,索邦大学)

专题命中 其他安全 :safety(abstract);分类 cs.CL、cs.AI

Comments ICCV 2025. The first three authors contributed equally. Project page and code: https://pegah- kh.github.io/projects/lmm-finetuning-analysis-and-steering/

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.09199 2025-08-14 cs.CV cs.AI cs.CL 62%

$Δ$-AttnMask: Attention-Guided Masked Hidden States for Efficient Data Selection and Augmentation

Jucheng Hu, Suorong Yang, Dongzhan Zhou

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.09919 2025-08-14 eess.IV cs.AI cs.CV 57%

T-CACE: A Time-Conditioned Autoregressive Contrast Enhancement Multi-Task Framework for Contrast-Free Liver MRI Synthesis, Segmentation, and Diagnosis

Xiaojiao Xiao, Jianfeng Zhao, Qinmin Vivian Hu, Guanghui Wang

机构 * Department of Computer Science, Toronto Metropolitan University(计算机科学系,多伦多 Metropolitan 大学) School of Biomedical Engineering, Western University(生物医学工程学院,西部大学)

专题命中 其他安全 :safety(abstract);分类 cs.AI

Comments IEEE Journal of Biomedical and Health Informatics, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.23077 2025-08-14 cs.CL 57%

Efficient Inference for Large Reasoning Models: A Survey

Yue Liu, Jiaying Wu, Yufei He, Ruihan Gong, Jun Xia, Liang Li, Hongcheng Gao, Hongyu Chen, Baolong Bi, Jiaheng Zhang, Zhiqi Huang, Bryan Hooi, Stan Z. Li, Keqin Li

专题命中 其他安全 :safety(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.09453 2025-08-14 cs.CV cs.LG 57%

HyperKD: Distilling Cross-Spectral Knowledge in Masked Autoencoders via Inverse Domain Shift with Spatial-Aware Masking and Specialized Loss

Abdul Matin, Tanjim Bin Faruk, Shrideep Pallickara, Sangmi Lee Pallickara

机构 * Colorado State University(科罗拉多州立大学)

专题命中 其他安全 :alignment(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏