arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-08-29 至 2025-08-29 共收录 10 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 其他安全 10 篇

2508.20130 2025-08-29 q-bio.QM cs.AI cs.LG 81%

Artificial Intelligence for CRISPR Guide RNA Design: Explainable Models and Off-Target Safety

Alireza Abbaszadeh, Armita Shahlai

机构 * Department of Computer Engineering, Ma.C., Islamic Azad University(计算机工程系,伊斯兰阿兹德大学) Department of Biological Sciences and Technologies, Faculty of Basic Sciences, Islamic Azad University(基础科学学院生物科学与技术系,伊斯兰阿兹德大学)

专题命中 其他安全 :safety(title,abstract);分类 cs.AI、cs.LG

Comments 29 pages, 5 figures, 2 tables, 42 cited references

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20465 2025-08-29 q-bio.NC 78%

On the possibility of deep alignment

Alex B. Kiefer

专题命中 其他安全 :alignment(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.19512 2025-08-29 cs.CL 70%

Safeguard Fine-Tuned LLMs Through Pre- and Post-Tuning Model Merging

Hua Farn, Hsuan Su, Shachi H Kumar, Saurav Sahay, Shang-Tse Chen, Hung-yi Lee

机构 * National Taiwan University(国立台湾大学) Intel Lab(英特尔实验室)

专题命中 其他安全 :alignment(abstract);safety(abstract);分类 cs.CL

Comments EMNLP 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20217 2025-08-29 cs.CL cs.AI 62%

Prompting Strategies for Language Model-Based Item Generation in K-12 Education: Bridging the Gap Between Small and Large Language Models

Mohammad Amini, Babak Ahmadi, Xiaomeng Xiong, Yilin Zhang, Christopher Qiao

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20722 2025-08-29 cs.CL 57%

rStar2-Agent: Agentic Reasoning Technical Report

Ning Shang, Yifei Liu, Yi Zhu, Li Lyna Zhang, Weijiang Xu, Xinyu Guan, Buze Zhang, Bingcheng Dong, Xudong Zhou, Bowen Zhang, Ying Xin, Ziming Miao, Scarlett Li, Fan Yang, Mao Yang

机构 * Microsoft Research(微软研究院)

专题命中 其他安全 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20374 2025-08-29 cs.AI 57%

TCIA: A Task-Centric Instruction Augmentation Method for Instruction Finetuning

Simin Ma, Shujian Liu, Jun Tan, Yebowen Hu, Song Wang, Sathish Reddy Indurthi, Sanqiang Zhao, Liwei Wu, Jianbing Han, Kaiqiang Song

专题命中 其他安全 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.15576 2025-08-29 cs.CV cs.LG 57%

Visual Perturbation and Adaptive Hard Negative Contrastive Learning for Compositional Reasoning in Vision-Language Models

Xin Huang, Ruibin Li, Tong Jia, Wei Zheng, Ya Wang

机构 * School of Artificial Intelligence and Software Engineering, Nanyang Normal University, Henan, China(人工智能与软件工程学院,南阳师范学院,河南) Institute for Artificial Intelligence, Peking University, Beijing, China(人工智能研究院,北京大学,北京) Collaborative Innovation Center of Intelligent Explosion-proof Equipment, Henan, China(智能防爆设备协同创新中心,河南)

专题命中 其他安全 :alignment(abstract);分类 cs.LG

Comments Accepted at the International Joint Conference on Artificial Intelligence (IJCAI 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.00258 2025-08-29 cs.AI q-bio.NC 57%

Possible Principles for Aligned Structure Learning Agents

Lancelot Da Costa, Tomáš Gavenčiak, David Hyland, Mandana Samiei, Cristian Dragos-Manta, Candice Pattisapu, Adeel Razi, Karl Friston

机构 * VERSES AI Research Lab(VERSES AI研究实验室) Charles University(查尔斯大学) University of Oxford(牛津大学) Mila, Quebec AI Institute(魁北克人工智能研究院) McGill University(麦吉尔大学) University of Montreal(蒙特利尔大学) University College London(伦敦大学学院) Monash University(莫纳什大学) CIFAR Azrieli Global Scholars Program(CIFAR阿兹里埃利全球学者计划)

专题命中 其他安全 :alignment(abstract);分类 cs.AI

Comments 24 pages of content, 33 with references; accepted version

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20623 2025-08-29 cs.CV 50%

AvatarBack: Back-Head Generation for Complete 3D Avatars from Front-View Images

Shiqi Xin, Xiaolin Zhang, Yanbin Liu, Peng Zhang, Caifeng Shan

机构 * College of Electrical Engineering and Automation, Shandong University of Science and Technology(山东科技大学电气工程与自动化学院) Department of Data Science and Artificial Intelligence, Auckland University of Technology(奥克兰大学数据科学与人工智能系) College of Computer Science and Engineering, Shandong University of Science and Technology(山东科技大学计算机科学与工程学院)

专题命中 其他安全 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14330 2025-08-29 cs.SE 50%

Leveraging LLMs for Formal Software Requirements -- Challenges and Prospects

Arshad Beg, Diarmuid O'Donoghue, Rosemary Monahan

专题命中 其他安全 :safety(abstract)

Comments Overlay2025 - 7th International Workshop on Artificial Intelligence and fOrmal VERification, Logic, Automata, and sYnthesis. [Accepted]. To be held on 26th of October, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏