arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-10-01 至 2025-10-01 共收录 9 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 其他安全 9 篇

2509.25253 2025-10-01 cs.LG cs.AI 81%

Knowledge distillation through geometry-aware representational alignment

Prajjwal Bhattarai, Mohammad Amjad, Dmytro Zhylko, Tuka Alhanai

机构 * New York University Abu Dhabi(纽约大学阿布扎赫德分校) New York University(纽约大学) Tandon School of Engineering(Tandon工程学院)

专题命中 其他安全 :alignment(title,abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.19607 2025-10-01 cs.HC cs.AI 79%

Enabling Rapid Shared Human-AI Mental Model Alignment via the After-Action Review

Edward Gu, Ho Chit Siu, Melanie Platt, Isabelle Hurley, Jaime Peña, Rohan Paleja

专题命中 其他安全 :alignment(title,abstract);分类 cs.AI

Comments Accepted to the Cooperative Multi-Agent Systems Decision-making and Learning:Human-Multi-Agent Cognitive Fusion Workshop at AAAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.26625 2025-10-01 cs.LG cs.AI cs.CV cs.MM 62%

Learning to See Before Seeing: Demystifying LLM Visual Priors from Language Pre-training

Junlin Han, Shengbang Tong, David Fan, Yufan Ren, Koustuv Sinha, Philip Torr, Filippos Kokkinos

机构 * Meta Superintelligence Labs(Meta 超智能实验室) University of Oxford(牛津大学)

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

Comments Project page: https://junlinhan.github.io/projects/lsbs/

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.26239 2025-10-01 cs.LG cs.AI stat.ML 62%

Sandbagging in a Simple Survival Bandit Problem

Joel Dyer, Daniel Jarne Ornia, Nicholas Bishop, Anisoara Calinescu, Michael Wooldridge

机构 * University of Oxford(牛津大学)

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG

Comments Forthcoming in the "Reliable ML from Unreliable Data Workshop" at NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25220 2025-10-01 cs.CL cs.LG 62%

Cyclic Ablation: Testing Concept Localization against Functional Regeneration in AI

Eduard Kapelko

机构 * Eduard Kapelko(独立研究者)

专题命中 其他安全 :safety(abstract);分类 cs.CL、cs.LG

Comments Code is available at: https://www.kaggle.com/code/kapedalex/cycleablationpublic/

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.26128 2025-10-01 cs.AI 57%

MEDAKA: Construction of Biomedical Knowledge Graphs Using Large Language Models

Asmita Sengupta, David Antony Selby, Sebastian Josef Vollmer, Gerrit Großmann

机构 * Department of Data Science and its Applications, German Research Center for Artificial Intelligence (DFKI GmbH)(数据科学及其应用系,德国人工智能研究中心(DFKI GmbH)) Department of Computer Science, University of Kaiserslautern–Landau (RPTU)(计算机科学系,凯撒斯劳滕-兰道大学(RPTU))

专题命中 其他安全 :safety(abstract);分类 cs.AI

Comments 9 pages, 5 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25243 2025-10-01 cs.SE cs.AI 57%

Reinforcement Learning-Guided Chain-of-Draft for Token-Efficient Code Generation

Xunzhu Tang, Iyiola Emmanuel Olatunji, Tiezhu Sun, Jacques Klein, Tegawende F. Bissyande

机构 * University of Luxembourg(卢森堡大学)

专题命中 其他安全 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.26012 2025-10-01 cs.CV 50%

SETR: A Two-Stage Semantic-Enhanced Framework for Zero-Shot Composed Image Retrieval

Yuqi Xiao, Yingying Zhu

机构 * Yuqi Xiao, Yingying Zhu

专题命中 其他安全 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25728 2025-10-01 cond-mat.mtrl-sci physics.app-ph 50%

Fingerprinting Organic Molecules for the Inverse Design of Two-Dimensional Hybrid Perovskites with Target Energetics

Yongxin Lyu, Yifan Zhou, Yu Zhang, Yang Yang, Bosen Zou, Qiang Weng, Tong Xie, Claudio Cazorla, Jianhua Hao, Jun Yin, Tom Wu

专题命中 其他安全 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏