arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-10-06 至 2025-10-06 共收录 5 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全训练 5 篇

2510.02768 2025-10-06 cs.LG cs.CL 81%

A Granular Study of Safety Pretraining under Model Abliteration

Shashank Agnihotri, Jonas Jakubassa, Priyam Dey, Sachin Goyal, Bernt Schiele, Venkatesh Babu Radhakrishnan, Margret Keuper

机构 * Data and Web Science Group, University of Mannheim(曼海姆大学数据与网络科学组) Vision and AI Lab, Indian Institute of Science(印度科学院视觉与人工智能实验室) Carnegie Mellon University(卡内基梅隆大学) Max-Planck-Institute for Informatics, Saarland Informatics Campus(马克斯·普朗克信息研究所,萨尔兰信息校园)

专题命中 安全训练 :safety(title,abstract);分类 cs.CL、cs.LG

Comments Accepted at NeurIPS 2025 bWorkshop Lock-LLM. *Equal Contribution

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02363 2025-10-06 eess.SY cs.SY 78%

Precise HDV Positioning through Safety-Aware Integrated Sensing and Communication in a Value-of-Information-Driven 6G V2X System

Mohammad Reza Abedi, Zahra Rashidi, Nader Mokari, Hamid Saeedi, Nizar Zorba

专题命中 安全训练 :safety(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.17746 2025-10-06 cs.LG cs.AI cs.CL 67%

Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains

Anisha Gunjal, Anthony Wang, Elaine Lau, Vaskar Nath, Yunzhong He, Bing Liu, Sean Hendryx

专题命中 安全训练 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02627 2025-10-06 cs.RO cs.AI 57%

A Trajectory Generator for High-Density Traffic and Diverse Agent-Interaction Scenarios

Ruining Yang, Yi Xu, Yixiao Chen, Yun Fu, Lili Su

机构 * Department of Electrical and Computer Engineering, Northeastern University(电气与计算机工程系,东北大学)

专题命中 安全训练 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02403 2025-10-06 q-bio.QM cs.AI cs.CV 57%

Glaucoma Detection and Structured OCT Report Generation via a Fine-tuned Multimodal Large Language Model

Jalil Jalili, Yashraj Gavhane, Evan Walker, Anna Heinke, Christopher Bowd, Akram Belghith, Massimo A. Fazio, Christopher A. Girkin, C. Gustavo De Moraes, Jeffrey M. Liebmann, Sally L. Baxter, Robert N. Weinreb, Linda M. Zangwill, Mark Christopher

专题命中 安全训练 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏