Joint Verification and Refinement of Language Models for Safety-Constrained Planning
机构 * University of Texas at Austin(德克萨斯大学奥斯汀分校)
专题命中 安全训练 :safety(title,abstract);分类 cs.AI
AI 大模型
大模型对齐、安全、越狱、红队、提示注入和可信评测。
机构 * University of Texas at Austin(德克萨斯大学奥斯汀分校)
专题命中 安全训练 :safety(title,abstract);分类 cs.AI
机构 * NVIDIA
专题命中 安全训练 :alignment(abstract);safety(abstract);分类 cs.CL、cs.AI、cs.LG
Comments Accepted at the Multi-Turn Interactions workshop at the 39th Conference on Neural Information Processing Systems (NeurIPS 2025)
机构 * School of Electronic and Electrical Engineering, Shanghai University of Engineering Science(上海工程技术大学电子与电气工程学院) ; Tencent Youtu Lab(腾讯优图实验室) ; Artificial Intelligence Innovation and Incubation Institute, Fudan University(复旦大学人工智能创新与孵化院)
专题命中 安全训练 :alignment(abstract);分类 cs.CL