arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-10-24 至 2025-10-24 共收录 8 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 其他安全 8 篇

2510.20377 2025-10-24 cs.AI cs.CL 62%

IKnow: Instruction-Knowledge-Aware Continual Pretraining for Effective Domain Adaptation

Tianyi Zhang, Florian Mai, Lucie Flek

机构 * University of Bonn(波恩大学) Lamarr Institute for Machine Learning and Artificial Intelligence(拉玛尔机器学习与人工智能研究所)

专题命中 其他安全 :safety(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.16722 2025-10-24 cs.CL cs.AI 62%

Breaking mBad! Supervised Fine-tuning for Cross-Lingual Detoxification

Himanshu Beniwal, Youngwoo Kim, Maarten Sap, Soham Dan, Thomas Hartvigsen

机构 * Indian Institute of Technology Gandhinagar(印度古吉拉特邦理工学院加尔文加尔)

专题命中 其他安全 :safety(abstract);分类 cs.CL、cs.AI

Comments Accepted at MELT Workshop @ COLM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20428 2025-10-24 cs.LG 57%

An Empirical Study of Sample Selection Strategies for Large Language Model Repair

Xuran Li, Jingyi Wang

机构 * Zhejiang University(浙江大学)

专题命中 其他安全 :safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20244 2025-10-24 cs.CV cs.LG 57%

Empower Words: DualGround for Structured Phrase and Sentence-Level Temporal Grounding

Minseok Kang, Minhyeok Lee, Minjung Kim, Donghyeong Kim, Sangyoun Lee

机构 * Yonsei University(延世大学) LG Electronics(LG电子)

专题命中 其他安全 :alignment(abstract);分类 cs.LG

Comments Comments: 28 pages, including appendix. 5 figures. Full version of the NeurIPS 2025 paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25033 2025-10-24 cs.CV cs.LG 57%

VT-FSL: Bridging Vision and Text with LLMs for Few-Shot Learning

Wenhao Li, Qiangchang Wang, Xianjing Meng, Zhibin Wu, Yilong Yin

机构 * School of Software, Shandong University(山东大学软件学院) Shenzhen Loop Area Institute(深圳河套学院) School of Computing and Artificial Intelligence, Shandong University of Finance and Economics(山东财经大学计算机与人工智能学院)

专题命中 其他安全 :alignment(abstract);分类 cs.LG

Comments Accepted by NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.19529 2025-10-24 cs.CL 57%

Blockwise SFT for Diffusion Language Models: Reconciling Bidirectional Attention and Autoregressive Decoding

Bowen Sun, Yujun Cai, Ming-Hsuan Yang, Yiwei Wang

机构 * University of California, Merced(加州大学默塞德分校) The University of Queensland(昆士兰大学) Google DeepMind(谷歌DeepMind)

专题命中 其他安全 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.19201 2025-10-24 cs.CL 57%

DREAM: Drafting with Refined Target Features and Entropy-Adaptive Cross-Attention Fusion for Multimodal Speculative Decoding

Yunhai Hu, Tianhua Xia, Zining Liu, Rahul Raman, Xingyu Liu, Bo Bao, Eric Sather, Vithursan Thangarasa, Sai Qian Zhang

机构 * Courant Institute of Mathematical Sciences, New York University(纽约大学数学科学学院) Tandon School of Engineering, New York University(纽约大学工程学院) Cerebras Systems Inc.(Cerebras Systems公司) University of Pennsylvania(宾夕法尼亚大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.14349 2025-10-24 cs.CV 50%

Vision-Centric Activation and Coordination for Multimodal Large Language Models

Yunnan Wang, Fan Lu, Kecheng Zheng, Ziyuan Huang, Ziqiang Li, Wenjun Zeng, Xin Jin

机构 * MoE Key Lab of Artificial Intelligence, Shanghai Jiao Tong University(人工智能MoE实验室,上海交通大学) Ant Group(蚂蚁集团) Ningbo Institute of Digital Twin, Eastern Institute of Technology, Ningbo(宁波数字孪生研究所,东部技术研究所,宁波)

专题命中 其他安全 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏