arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-09-16 至 2025-09-16 共收录 12 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 其他安全 12 篇

2410.15633 2025-09-16 cs.CL cs.AI 81%

GATEAU: Selecting Influential Samples for Long Context Alignment

Shuzheng Si, Haozhe Zhao, Gang Chen, Yunshui Li, Kangyang Luo, Chuancheng Lv, Kaikai An, Fanchao Qi, Baobao Chang, Maosong Sun

机构 * Department of Computer Science and Technology, Tsinghua University(清华大学计算机科学与技术系) Peking University(北京大学) DeepLang AI Institute for AI, Tsinghua University(清华大学人工智能研究院)

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI

Comments EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.09931 2025-09-16 cs.LG cs.AI 81%

Mechanistic Interpretability of LoRA-Adapted Language Models for Nuclear Reactor Safety Applications

Yoon Pyo Lee

专题命中 其他安全 :safety(title,abstract);分类 cs.AI、cs.LG

Comments Accepted for publication in Nuclear Technology. 24 pages, 2 tables, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10509 2025-09-16 cs.LG cs.AI 73%

The Anti-Ouroboros Effect: Emergent Resilience in Large Language Models from Recursive Selective Feedback

Sai Teja Reddy Adapala

机构 * University of North Carolina at Charlotte(北卡罗来纳大学夏洛特分校)

专题命中 其他安全 :safety(abstract);AI safety(abstract);分类 cs.AI、cs.LG

Comments 5 pages, 3 figures, 2 tables. Code is available at: https://github.com/imsaitejareddy/ouroboros-effect-experiment

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10529 2025-09-16 cs.LG cs.AI cs.CV 62%

Mitigating Catastrophic Forgetting and Mode Collapse in Text-to-Image Diffusion via Latent Replay

Aoi Otani

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10523 2025-09-16 cs.LG cs.AI 62%

From Predictions to Explanations: Explainable AI for Autism Diagnosis and Identification of Critical Brain Regions

Kush Gupta, Amir Aly, Emmanuel Ifeachor, Rohit Shankar

机构 * University of Plymouth, Plymouth, UK(普利茅斯大学)

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.01042 2025-09-16 cs.LG 57%

SafeSwitch: Steering Unsafe LLM Behavior via Internal Activation Signals

Peixuan Han, Cheng Qian, Xiusi Chen, Yuji Zhang, Heng Ji, Denghui Zhang

机构 * Siebel School of Computing(计算科学系) Data Science, University of Illinois Urbana-Champaign(数据科学,伊利诺伊大学厄巴纳-香槟分校) School of Business, Stevens Institute of Technology(商学院,史蒂文斯理工学院)

专题命中 其他安全 :safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11661 2025-09-16 cs.CV cs.AI 57%

DTGen: Generative Diffusion-Based Few-Shot Data Augmentation for Fine-Grained Dirty Tableware Recognition

Lifei Hao, Yue Cheng, Baoqi Huang, Bing Jia, Xuandong Zhao

专题命中 其他安全 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.00521 2025-09-16 cs.SE cs.AI 57%

Automated detection of atomicity violations in large-scale systems

Hang He, Yixing Luo, Chengcheng Wan, Ting Su, Haiying Sun, Geguang Pu

机构 * East China Normal University(华东师范大学) Beijing Institute of Control Engineering(北京控制工程研究所) East China Normal University Shanghai Innovation Institute(华东师范大学上海创新研究院)

专题命中 其他安全 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11840 2025-09-16 cs.CV 50%

Synthetic Captions for Open-Vocabulary Zero-Shot Segmentation

Tim Lebailly, Vijay Veerabadran, Satwik Kottur, Karl Ridgeway, Michael Louis Iuzzolino

机构 * Meta KU Leuven(鲁汶大学)

专题命中 其他安全 :alignment(abstract)

Comments ICCV 2025 CDEL Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.12356 2025-09-16 eess.IV cs.CV 50%

Regist3R: Incremental Registration with Stereo Foundation Model

Sidun Liu, Wenyu Li, Peng Qiao, Yong Dou

机构 * College of Computer Science and Technology(计算机科学与技术学院) National Key Laboratory of Parallel and Distributed Computing(并行与分布式计算国家重点实验室) National University of Defense Technology(国防科技大学)

专题命中 其他安全 :alignment(abstract)

Comments Accepted by ACM Multimedia 2025. github link: https://github.com/Liu-SD/Regist3R

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11783 2025-09-16 cs.RO 50%

Augmented Reality-Enhanced Robot Teleoperation for Collecting User Demonstrations

Shiqi Gong, Sebastian Zudaire, Chi Zhang, Zhen Li

机构 * Aalto University(阿alto大学) ABB Corporate Research(ABB企业研究)

专题命中 其他安全 :safety(abstract)

Comments Accepted by 2025 8th International Conference on Robotics, Control and Automation Engineering (RCAE 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10767 2025-09-16 cs.CV 50%

Enhancement Without Contrast: Stability-Aware Multicenter Machine Learning for Glioma MRI Imaging

Sajad Amiri, Shahram Taeb, Sara Gharibi, Setareh Dehghanfard, Somayeh Sadat Mehrnia, Mehrdad Oveisi, Ilker Hacihaliloglu, Arman Rahmim, Mohammad R. Salmanpour

专题命中 其他安全 :safety(abstract)

Comments 14 Pages, 1 Figure, and 6 Tables

详情

展开后加载摘要…

URL PDF HTML 收藏