arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 8034 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 其他安全 8034 篇

2601.03047 2026-01-07 cs.LG 70%

When the Coffee Feature Activates on Coffins: An Analysis of Feature Extraction and Steering for Mechanistic Interpretability

当咖啡特征在棺材上激活:对特征提取和转向用于机制可解释性的分析

Raphael Ronge, Markus Maier, Frederick Eberhardt

机构 * Department of Philosophy of Nature and Technology(自然哲学与技术系) Munich School of Philosophy(慕尼黑哲学学院) Division of the Humanities and Social Sciences(人文与社会科学系) California Institute of Technology(加州理工学院)

专题命中 其他安全 :safety(abstract);AI safety(abstract);分类 cs.LG

AI总结 本文分析了通过稀疏自编码器提取特征和控制模型输出的方法,指出其在机制可解释性中的局限性和可靠性问题,强调需转向更可靠的预测与控制。

Comments 33 pages (65 with appendix), 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.14500 2025-11-04 cs.AI cs.MA cs.NE 70%

The Digital Ecosystem of Beliefs: does evolution favour AI over humans?

David M. Bossens, Shanshan Feng, Yew-Soon Ong

机构 * Institute of High Performance Computing (IHPC), Agency for Science, Technology and Research (A*STAR) Centre for Frontier AI Research (CFAR), Agency for Science, Technology and Research (A*STAR)(高性能计算研究所(IHPC)、科技研究局(A*STAR)前沿人工智能研究中心(CFAR)、科技研究局(A*STAR)) School of Computer Science Wuhan University(计算机科学学院 武汉大学)

专题命中 其他安全 :safety(abstract);AI safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11040 2025-10-14 cs.CL 70%

Enabling Doctor-Centric Medical AI with LLMs through Workflow-Aligned Tasks and Benchmarks

Wenya Xie, Qingying Xiao, Yu Zheng, Xidong Wang, Junying Chen, Ke Ji, Anningzhe Gao, Prayag Tiwari, Xiang Wan, Feng Jiang, Benyou Wang

机构 * Shenzhen Research Institute of Big Data(大数据研究 institute) National Health Data Institute(国家健康数据研究所) Halmstad University(哈马碧大学) Shenzhen University of Advanced Technology(深圳先进技术大学)

专题命中 其他安全 :alignment(abstract);safety(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01363 2025-10-03 cs.AI 70%

Retrieval-Augmented Framework for LLM-Based Clinical Decision Support

Leon Garza, Anantaa Kotal, Michael A. Grasso, Emre Umucu

机构 * Dept. of Computer Science, The University of Texas at El Paso, USA(计算机科学系,德克萨斯大学埃尔帕索分校) Dept. of Emergency Medicine, University of Maryland School of Medicine, USA(急诊医学系,马里兰大学医学院) Dept. of Public Health Sciences, The University of Texas at El Paso, USA(公共卫生科学系,德克萨斯大学埃尔帕索分校)

专题命中 其他安全 :alignment(abstract);safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.00300 2025-10-02 cs.AI 70%

ICL Optimized Fragility

Serena Gomez Wannaz

机构 * Serena Gomez Wannaz(独立研究者)

专题命中 其他安全 :safety(abstract);AI safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16660 2025-09-23 cs.CL 70%

Redefining Experts: Interpretable Decomposition of Language Models for Toxicity Mitigation

Zuhair Hasan Shaik, Abdullah Mazhar, Aseem Srivastava, Md Shad Akhtar

机构 * IIIT Dharwad, India(印度IIIT达尔瓦德大学) IIIT Delhi, India(印度IIIT德里大学) FLaME-NLP Lab, IIIT Delhi(IIIT德里大学FLaME-NLP实验室)

专题命中 其他安全 :safety(abstract);AI safety(abstract);分类 cs.CL

Comments Accepted to the NeurIPS 2025 Research Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.19512 2025-08-29 cs.CL 70%

Safeguard Fine-Tuned LLMs Through Pre- and Post-Tuning Model Merging

Hua Farn, Hsuan Su, Shachi H Kumar, Saurav Sahay, Shang-Tse Chen, Hung-yi Lee

机构 * National Taiwan University(国立台湾大学) Intel Lab(英特尔实验室)

专题命中 其他安全 :alignment(abstract);safety(abstract);分类 cs.CL

Comments EMNLP 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15796 2025-08-25 cs.CL cs.AI cs.CY cs.LG 70%

Benchmarking the Legal Reasoning of LLMs in Arabic Islamic Inheritance Cases

Nouar AlDahoul, Yasir Zaki

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.CY

Comments 5 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01059 2025-08-05 cs.CR cs.AI 70%

Llama-3.1-FoundationAI-SecurityLLM-8B-Instruct Technical Report

Sajana Weerawardhena, Paul Kassianik, Blaine Nelson, Baturay Saglam, Anu Vellore, Aman Priyanshu, Supriti Vijay, Massimo Aufiero, Arthur Goldblatt, Fraser Burch, Ed Li, Jianliang He, Dhruv Kedia, Kojin Oshiba, Zhouran Yang, Yaron Singer, Amin Karbasi

专题命中 其他安全 :alignment(abstract);safety(abstract);分类 cs.AI

Comments 34 pages - Technical Report

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20028 2025-07-29 cs.CV cs.AI 70%

TAPS : Frustratingly Simple Test Time Active Learning for VLMs

Dhruv Sarkar, Aprameyo Chakrabartty, Bibhudatta Bhanja

机构 * IIT Kharagpur(印度理工学院Kharagpur分校)

专题命中 其他安全 :alignment(abstract);safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.21892 2025-06-30 cs.CV cs.AI 70%

SODA: Out-of-Distribution Detection in Domain-Shifted Point Clouds via Neighborhood Propagation

Adam Goodge, Xun Xu, Bryan Hooi, Wee Siong Ng, Jingyi Liao, Yongyi Su, Xulei Yang

机构 * Institute for Infocomm Research, Agency for Science, Technology and Research (A*STAR), Singapore(信息通信研究所,科技研究局(A*STAR),新加坡) School of Computing, National University of Singapore, Singapore(计算学院,新加坡国立大学,新加坡)

专题命中 其他安全 :alignment(abstract);safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.17824 2025-06-24 quant-ph cs.CR cs.LG 70%

Quantum-Hybrid Support Vector Machines for Anomaly Detection in Industrial Control Systems

Tyler Cultice, Md. Saif Hassan Onim, Annarita Giani, Himanshu Thapliyal

机构 * Department of Electrical Engineering and Compute Science, University of Tennessee, Knoxville, TN, 37996 USA(电气工程与计算科学系,田纳西大学,诺克斯维尔,TN,37996 USA) GE Vernova Research Center(通用电气研究中心)

专题命中 其他安全 :alignment(abstract);safety(abstract);分类 cs.LG

Comments 12 pages, 6 tables, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.13419 2025-05-21 cs.CL cs.AI cs.CY cs.LG cs.SC 70%

From Words to Worlds: Compositionality for Cognitive Architectures

Ruchira Dhar, Anders Søgaard

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.CY

Comments Accepted to ICML 2024 Workshop on LLMs & Cognition

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.05501 2025-05-12 cs.CV cs.AI eess.IV 70%

Preliminary Explorations with GPT-4o(mni) Native Image Generation

Pu Cao, Feng Zhou, Junyi Ji, Qingye Kong, Zhixiang Lv, Mingjian Zhang, Xuekun Zhao, Siqi Wu, Yinghui Lin, Qing Song, Lu Yang

专题命中 其他安全 :alignment(abstract);safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.04070 2025-04-08 cs.MA cs.AI 70%

Enforcement Agents: Enhancing Accountability and Resilience in Multi-Agent AI Frameworks

Sagar Tamang, Dibya Jyoti Bora

专题命中 其他安全 :alignment(abstract);safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.01850 2025-04-03 cs.SE cs.AI 70%

Code Red! On the Harmfulness of Applying Off-the-shelf Large Language Models to Programming Tasks

Ali Al-Kaswan, Sebastian Deatc, Begüm Koç, Arie van Deursen, Maliheh Izadi

专题命中 其他安全 :alignment(abstract);harmlessness(abstract);分类 cs.AI

Comments FSE'25 Technical Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.09555 2025-03-04 cs.CV cs.AI 70%

Text-driven Adaptation of Foundation Models for Few-shot Surgical Workflow Analysis

Tingxuan Chen, Kun Yuan, Vinkle Srivastav, Nassir Navab, Nicolas Padoy

专题命中 其他安全 :alignment(abstract);safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.23496 2025-03-04 cs.CL 70%

Smaller Large Language Models Can Do Moral Self-Correction

Guangliang Liu, Zhiyu Xue, Xitong Zhang, Rongrong Wang, Kristen Marie Johnson

专题命中 其他安全 :alignment(abstract);safety(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.07587 2025-02-12 cs.LG 70%

SEMU: Singular Value Decomposition for Efficient Machine Unlearning

Marcin Sendera, Łukasz Struski, Kamil Książek, Kryspin Musiol, Jacek Tabor, Dawid Rymarczyk

专题命中 其他安全 :safety(abstract);AI safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.18438 2025-02-03 cs.SE cs.AI 70%

o3-mini vs DeepSeek-R1: Which One is Safer?

Aitor Arrieta, Miriam Ugarte, Pablo Valle, José Antonio Parejo, Sergio Segura

专题命中 其他安全 :alignment(abstract);safety(abstract);分类 cs.AI

Comments arXiv admin note: substantial text overlap with arXiv:2501.17749

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.16513 2025-01-31 cs.CL 70%

Deception in LLMs: Self-Preservation and Autonomous Goals in Large Language Models

Sudarshan Kamath Barkur, Sigurd Schacht, Johannes Scholl

专题命中 其他安全 :alignment(abstract);safety(abstract);分类 cs.CL

Comments Corrected Version - Solved Some Issues with reference compilation by latex

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.06167 2025-01-08 cs.AI 70%

Predictable Artificial Intelligence

Lexin Zhou, Pablo A. Moreno-Casares, Fernando Martínez-Plumed, John Burden, Ryan Burnell, Lucy Cheke, Cèsar Ferri, Alexandru Marcoci, Behzad Mehrbakhsh, Yael Moros-Daval, Seán Ó hÉigeartaigh, Danaja Rutar, Wout Schellaert, Konstantinos Voudouris, José Hernández-Orallo

专题命中 其他安全 :alignment(abstract);safety(abstract);分类 cs.AI

Comments Paper Under Review

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.05285 2024-12-03 cs.AI cs.SE 70%

AgentOps: Enabling Observability of LLM Agents

Liming Dong, Qinghua Lu, Liming Zhu

专题命中 其他安全 :safety(abstract);AI safety(abstract);分类 cs.AI

Comments 12 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.13944 2024-10-21 cs.CL 70%

Boosting LLM Translation Skills without General Ability Loss via Rationale Distillation

Junhong Wu, Yang Zhao, Yangyifan Xu, Bing Liu, Chengqing Zong

专题命中 其他安全 :alignment(abstract);safety(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.17287 2024-10-01 cs.CL 70%

When to Trust LLMs: Aligning Confidence with Response Quality

Shuchang Tao, Liuyi Yao, Hanxing Ding, Yuexiang Xie, Qi Cao, Fei Sun, Jinyang Gao, Huawei Shen, Bolin Ding

专题命中 其他安全 :alignment(abstract);safety(abstract);分类 cs.CL

Comments Accepted by ACL 2024. Code: https://github.com/TaoShuchang/CONQORD

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.04792 2024-07-12 cs.GT cs.AI 70%

Playing Large Games with Oracles and AI Debate

Xinyi Chen, Angelica Chen, Dean Foster, Elad Hazan

专题命中 其他安全 :safety(abstract);AI safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.03820 2024-06-24 cs.CL 70%

CantTalkAboutThis: Aligning Language Models to Stay on Topic in Dialogues

Makesh Narsimhan Sreedhar, Traian Rebedea, Shaona Ghosh, Jiaqi Zeng, Christopher Parisien

专题命中 其他安全 :alignment(abstract);safety(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.13669 2024-05-29 cs.CL 70%

Self-Distillation Bridges Distribution Gap in Language Model Fine-Tuning

Zhaorui Yang, Tianyu Pang, Haozhe Feng, Han Wang, Wei Chen, Minfeng Zhu, Qian Liu

专题命中 其他安全 :alignment(abstract);safety(abstract);分类 cs.CL

Comments ACL 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.16967 2024-04-29 cs.LG cs.CR 70%

ML2SC: Deploying Machine Learning Models as Smart Contracts on the Blockchain

Zhikai Li, Steve Vott, Bhaskar Krishnamachar

专题命中 其他安全 :safety(abstract);AI safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.10636 2024-04-18 cs.CY cs.AI cs.CL cs.HC cs.LG 70%

What are human values, and how do we align AI to them?

Oliver Klingefjord, Ryan Lowe, Joe Edelman

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏