arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-07-30 至 2025-07-30 共收录 42 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 16 篇

2507.21157 2025-07-30 cs.CR cs.CV 50%

Unmasking Synthetic Realities in Generative AI: A Comprehensive Review of Adversarially Robust Deepfake Detection Systems

Naseem Khan, Tuan Nguyen, Amine Bermak, Issa Khalil

机构 * Department of Computer Science(计算机科学系) Hamad bin Khalifa University(哈马德·本·哈利法大学) Qatar Computing Research Institute(卡塔尔计算研究所)

专题命中 安全评测 :trustworthy(abstract)

Comments 27 pages, 4 Tables, 3 Figures

详情

展开后加载摘要…

URL PDF HTML 收藏

2. AI治理与伦理 4 篇

2507.21091 2025-07-30 cs.CY cs.AI 81%

The Value of Gen-AI Conversations: A bottom-up Framework for AI Value Alignment

Lenart Motnikar, Katharina Baum, Alexander Kagan, Sarah Spiekermann-Hoff

专题命中 AI治理与伦理 :alignment(title,abstract);分类 cs.AI、cs.CY

Comments Thirty-Third European Conference on Information Systems (ECIS 2025), Amman, Jordan

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21083 2025-07-30 cs.CL cs.AI 62%

ChatGPT Reads Your Tone and Responds Accordingly -- Until It Does Not -- Emotional Framing Induces Bias in LLM Outputs

Franck Bardol

机构 * Independent Researcher(独立研究者)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21929 2025-07-30 cs.AI 57%

Libra: Large Chinese-based Safeguard for AI Content

Ziyang Chen, Huimu Yu, Xing Wu, Dongqin Liu, Songlin Hu

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21319 2025-07-30 cs.CL 57%

Do Large Language Models Understand Morality Across Cultures?

Hadi Mohammadi, Yasmeen F. S. S. Meijer, Efthymia Papadopoulou, Ayoub Bagheri

机构 * Department of Methodology and Statistics, Utrecht University, The Netherlands(方法论与统计学系,乌特雷赫特大学,荷兰)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 其他安全 7 篇

2507.21082 2025-07-30 cs.CY 79%

Safety Features for a Centralised AGI Project

Sarah Hastings-Woodhouse

专题命中 其他安全 :safety(title,abstract);分类 cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21839 2025-07-30 cs.CY cs.AI 73%

Against racing to AGI: Cooperation, deterrence, and catastrophic risks

Leonard Dung, Max Hellrigel-Holderbaum

专题命中 其他安全 :safety(abstract);AI safety(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21138 2025-07-30 cs.CL cs.AI cs.LG cs.SD eess.AS 67%

TTS-1 Technical Report

Oleg Atamanenko, Anna Chalova, Joseph Coombes, Nikki Cope, Phillip Dang, Zhifeng Deng, Jimmy Du, Michael Ermolenko, Feifan Fan, Yufei Feng, Cheryl Fichter, Pavel Filimonov, Louis Fischer, Kylan Gibbs, Valeria Gusarova, Pavel Karpik, Andreas Assad Kottner, Ian Lee, Oliver Louie, Jasmine Mai, Mikhail Mamontov, Suri Mao, Nurullah Morshed, Igor Poletaev, Florin Radu, Dmytro Semernia, Evgenii Shingarev, Vikram Sivaraja, Peter Skirko, Rinat Takhautdinov, Robert Villahermosa, Jean Wang

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments 20 pages, 10 figures. For associated modeling and training code, see https://github.com/inworld-ai/tts

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21107 2025-07-30 cs.CL cs.AI 62%

Curved Inference: Concern-Sensitive Geometry in Large Language Model Residual Streams

Rob Manson

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

Comments 29 pages, 22 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20536 2025-07-30 cs.CV cs.AI cs.HC 57%

T2I-Copilot: A Training-Free Multi-Agent Text-to-Image System for Enhanced Prompt Interpretation and Interactive Generation

Chieh-Yun Chen, Min Shi, Gong Zhang, Humphrey Shi

专题命中 其他安全 :alignment(abstract);分类 cs.AI

Comments ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21941 2025-07-30 eess.SY cs.SY 50%

Hierarchical Game-Based Multi-Agent Decision-Making for Autonomous Vehicles

Mushuang Liu, Yan Wan, Frank Lewis, Subramanya Nageshrao, H. Eric Tseng, Dimitar Filev

专题命中 其他安全 :safety(abstract)

Comments 12 pages, 20 figures, 1 algorithm

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21649 2025-07-30 cs.CV 50%

The Evolution of Video Anomaly Detection: A Unified Framework from DNN to MLLM

Shibo Gao, Peipei Yang, Haiyang Guo, Yangyang Liu, Yi Chen, Shuai Li, Han Zhu, Jian Xu, Xu-Yao Zhang, Linlin Huang

机构 * Beijing Jiaotong University(北京交通大学) State Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, Chinese Academy of Sciences(多模态人工智能系统国家重点实验室,自动化研究所,中国科学院) School of Artificial Intelligence, University of Chinese Academy of Sciences(人工智能学院,中国科学院大学) Zhongguancun Academy, Beijing, China(中关村学院,北京,中国)

专题命中 其他安全 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏