arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-10-27 至 2025-10-27 共收录 42 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. AI治理与伦理 4 篇

2510.21203 2025-10-27 cs.CY 57%

The Nuclear Analogy in AI Governance Research

Sophia Hatz

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CY

Comments Hatz, S. (in press). The Nuclear Analogy in AI Governance Research. In M. Furendal & M. Lundgren (Eds.), Handbook on the Global Governance of Artificial Intelligence. Edward Elgar Publishing

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.02992 2025-10-27 cs.AI cs.CL cs.LG 56%

Mitigating Manipulation and Enhancing Persuasion: A Reflective Multi-Agent Approach for Legal Argument Generation

Li Zhang, Kevin D. Ashley

机构 * Intelligent Systems Program University of Pittsburgh Pittsburgh Pennsylvania USA Intelligent Systems Program University of Pittsburgh

专题命中 AI治理与伦理 :分类 cs.CL、cs.AI、cs.LG;safety(comments)

Comments 13 pages, 2 figures, 2nd ConventicLe on Artificial Intelligence Regulation and Safety Workshop at ICAIL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 其他安全 10 篇

2510.21520 2025-10-27 cs.CL 79%

Brain-tuning Improves Generalizability and Efficiency of Brain Alignment in Speech Models

Omer Moussa, Mariya Toneva

机构 * Max Planck Institute for Software Systems(马克斯·普朗克软件系统研究所)

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL

Comments Published at the 39th Conference on Neural Information Processing Systems (NeurIPS 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.11194 2025-10-27 cs.CE 78%

Prot2Text-V2: Protein Function Prediction with Multimodal Contrastive Alignment

Xiao Fei, Michail Chatzianastasis, Sarah Almeida Carneiro, Hadi Abdine, Lawrence P. Petalidis, Michalis Vazirgiannis

专题命中 其他安全 :alignment(title,abstract)

Comments 24 pages, 11 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21184 2025-10-27 cs.LG cs.AI cs.CL stat.ML 67%

Reducing the Probability of Undesirable Outputs in Language Models Using Probabilistic Inference

Stephen Zhao, Aidan Li, Rob Brekelmans, Roger Grosse

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21623 2025-10-27 cs.CL cs.AI 62%

The Universal Landscape of Human Reasoning

Qiguang Chen, Jinhao Liu, Libo Qin, Yimeng Zhang, Yihao Liang, Shangxu Ren, Chengyu Luan, Dengyun Peng, Hanjing Li, Jiannan Guan, Zheng Yan, Jiaqi Wang, Mengkang Hu, Yantao Du, Zhi Chen, Xie Chen, Wanxiang Che

机构 * Harbin Institute of Technology(哈尔滨工业大学) Central South University(中南大学) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Princeton University(普林斯顿大学) The Chinese University of Hong Kong(香港中文大学) The University of Hong Kong(香港大学) ByteDance Seed (China)(字节跳动种子(中国)) Shanghai Jiao Tong University(上海交通大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

Comments Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21359 2025-10-27 cs.CL cs.AI 62%

Influence Guided Context Selection for Effective Retrieval-Augmented Generation

Jiale Deng, Yanyan Shen, Ziyuan Pei, Youmin Chen, Linpeng Huang

机构 * Shanghai Jiao Tong University(上海交通大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21501 2025-10-27 cs.CV cs.AI 57%

GranViT: A Fine-Grained Vision Model With Autoregressive Perception For MLLMs

Guanghao Zheng, Bowen Shi, Mingxing Xu, Ruoyu Sun, Peisen Zhao, Zhibo Zhang, Wenrui Dai, Junni Zou, Hongkai Xiong, Xiaopeng Zhang, Qi Tian

机构 * Shanghai Jiao Tong University(上海交通大学) Huawei Inc.(华为公司)

专题命中 其他安全 :alignment(abstract);分类 cs.AI

Comments 21 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.09396 2025-10-27 cs.AI cs.MA 57%

The Influence of Human-inspired Agentic Sophistication in LLM-driven Strategic Reasoners

Vince Trencsenyi, Agnieszka Mensfelt, Kostas Stathis

机构 * Department of Computer Science, Royal Holloway University of London(皇家霍洛威大学计算机科学系)

专题命中 其他安全 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21581 2025-10-27 cs.CV cs.SD 50%

Foley Control: Aligning a Frozen Latent Text-to-Audio Model to Video

Ciara Rowles, Varun Jampani, Simon Donné, Shimon Vainer, Julian Parker, Zach Evans

机构 * Stability AI

专题命中 其他安全 :alignment(abstract)

Comments Project Page: https://stability-ai.github.io/foleycontrol.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.16371 2025-10-27 cs.CV 50%

AGC-Drive: A Large-Scale Dataset for Real-World Aerial-Ground Collaboration in Driving Scenarios

Yunhao Hou, Bochao Zou, Min Zhang, Ran Chen, Shangdong Yang, Yanmei Zhang, Junbao Zhuo, Siheng Chen, Jiansheng Chen, Huimin Ma

机构 * University of Science and Technology Beijing(北京科技大学) Xiamen NEVC Advanced Electric Powertrain Technology Innovation Center(厦门NEVC先进电驱技术创新中心) Shanghai Jiao Tong University(上海交通大学)

专题命中 其他安全 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20859 2025-10-27 q-bio.NC 50%

Vision-language models learn the geometry of human perceptual space

Craig Sanders, Billy Dickson, Sahaj Singh Maini, Robert Nosofsky, Zoran Tiganj

专题命中 其他安全 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏