arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-08-28 至 2025-08-28 共收录 36 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 偏好对齐 3 篇

2508.19532 2025-08-28 cs.CL 79%

Alignment with Fill-In-the-Middle for Enhancing Code Generation

Houxing Ren, Zimu Lu, Weikang Shi, Haotian Hou, Yunqiao Yang, Ke Wang, Aojun Zhou, Junting Pan, Mingjie Zhan, Hongsheng Li

机构 * CUHK MMLab(CUHK语音实验室) SenseTime Research(商汤科技研究院) CPII under InnoHK(创新香港下的CPII) Shanghai AI Laboratory(上海人工智能实验室) Beihang University(北航)

专题命中 偏好对齐 :alignment(title);DPO(abstract);分类 cs.CL

Comments Accepted to EMNLP 2025 (main conference)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.19922 2025-08-28 cs.CL 70%

HEAL: A Hypothesis-Based Preference-Aware Analysis Framework

Yifu Huo, Chenglong Wang, Qiren Zhu, Shunjie Xing, Tong Xiao, Chunliang Zhang, Tongran Liu, Jinbo Zhu

机构 * School of Computer Science and Engineering, Northeastern University(计算机科学与工程学院,东北大学) NiuTrans Research(NiuTrans研究院) CAS Key Laboratory of Behavioral Science, Institute of Psychology, CAS(行为科学重点实验室,心理学研究所)

专题命中 偏好对齐 :alignment(abstract);DPO(abstract);分类 cs.CL

Comments Accepted by EMNLP 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.19567 2025-08-28 cs.LG 57%

Counterfactual Reward Model Training for Bias Mitigation in Multimodal Reinforcement Learning

Sheryl Mathew, N Harshit

专题命中 偏好对齐 :RLHF(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 安全训练 4 篇

2508.19562 2025-08-28 cs.AI 79%

Democracy-in-Silico: Institutional Design as Alignment in AI-Governed Polities

Trisanth Srinivasan, Santosh Patapati

机构 * Cyrion Labs(Cyron实验室)

专题命中 安全训练 :alignment(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.04867 2025-08-28 cs.AI cs.LG 66%

Think Smart, Act SMARL! Analyzing Probabilistic Logic Shields for Multi-Agent Reinforcement Learning

Satchit Chatterji, Erman Acar

机构 * IvI \& ILLC, University of Amsterdam

专题命中 安全训练 :safety(abstract,comments);分类 cs.AI、cs.LG

Comments Accepted to the 28th European Conference on Artificial Intelligence (ECAI 2025) --- 21 pages, 15 figures, Earlier title: "Analyzing Probabilistic Logic Driven Safety in Multi-Agent Reinforcement Learning"; (changed for specificity and clarity)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20040 2025-08-28 cs.AI cs.LG 62%

Model Science: getting serious about verification, explanation and control of AI systems

Przemyslaw Biecek, Wojciech Samek

机构 * Centre for Credible AI, Warsaw University of Technology, University of Warsaw(可信AI研究中心,华沙技术大学,华沙大学) Department of Artificial Intelligence, Fraunhofer Heinrich Hertz Institute, Technical University of Berlin, Germany BIFOLD - Berlin Institute for the Foundations of Learning(人工智能系,弗劳恩霍夫 Heinrich Hertz 研究所,柏林技术大学,德国 BIFOLD - 柏林学习与数据基础研究所)

专题命中 安全训练 :alignment(abstract);分类 cs.AI、cs.LG

Comments 8 pages

Journal ref Frontiers in AI track at European Conference on Artificial Intelligence (ECAI) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.19500 2025-08-28 cs.CR cs.AI 57%

Servant, Stalker, Predator: How An Honest, Helpful, And Harmless (3H) Agent Unlocks Adversarial Skills

David Noever

专题命中 安全训练 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 越狱攻击 3 篇

2508.19697 2025-08-28 cs.CR cs.AI cs.CL 90%

Safety Alignment Should Be Made More Than Just A Few Attention Heads

Chao Huang, Zefeng Zhang, Juewei Yue, Quangang Li, Chuang Zhang, Tingwen Liu

机构 * Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所) School of Cyber Security, University of Chinese Academy of Sciences(中国科学院大学网络安全学院)

专题命中 越狱攻击 :alignment(title,abstract);safety(title,abstract);jailbreak(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.19292 2025-08-28 cs.CR cs.AI 70%

Stand on The Shoulders of Giants: Building JailExpert from Previous Attack Experience

Xi Wang, Songlei Jian, Shasha Li, Xiaopeng Li, Bin Ji, Jun Ma, Xiaodong Liu, Jing Wang, Feilong Bao, Jianfeng Zhang, Baosheng Wang, Jie Yu

机构 * National University of Defense and Technology(国防科技大学) Inner Mongolia University(内蒙古大学)

专题命中 越狱攻击 :safety(abstract);jailbreak(abstract);分类 cs.AI

Comments 18 pages, EMNLP 2025 Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.19288 2025-08-28 cs.CR cs.AI 57%

Tricking LLM-Based NPCs into Spilling Secrets

Kyohei Shiomi, Zhuotao Lian, Toru Nakanishi, Teruaki Kitasuka

机构 * Hiroshima University(广岛大学)

专题命中 越狱攻击 :prompt injection(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 红队测试 1 篇

2508.19461 2025-08-28 cs.AI cs.CR cs.LG 73%

Reliable Weak-to-Strong Monitoring of LLM Agents

Neil Kale, Chen Bo Calvin Zhang, Kevin Zhu, Ankit Aich, Paula Rodriguez, Scale Red Team, Christina Q. Knight, Zifan Wang

机构 * Scale AI Carnegie Mellon University(卡内基梅隆大学) Massachusetts Institute of Technology(麻省理工学院)

专题命中 红队测试 :red teaming(abstract);prompt injection(abstract);分类 cs.AI、cs.LG

Comments 18 pages, 15 figures

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 幻觉与事实性 1 篇

2508.19432 2025-08-28 cs.AI 57%

Quantized but Deceptive? A Multi-Dimensional Truthfulness Evaluation of Quantized LLMs

Yao Fu, Xianxuan Long, Runchao Li, Haotian Yu, Mu Sheng, Xiaotian Han, Yu Yin, Pan Li

机构 * Case Western Reserve University(凯斯西储大学) Hangzhou Dianzi University(杭州电子科技大学)

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.AI

Comments Accepted to EMNLP2025 main conference (poster)

详情

展开后加载摘要…

URL PDF HTML 收藏

6. 隐私与版权 2 篇

2508.19495 2025-08-28 cs.DC cs.LG eess.SP 57%

Towards 6G Intelligence: The Role of Generative AI in Future Wireless Networks

Muhammad Ahmed Mohsin, Junaid Ahmad, Muhammad Hamza Nawaz, Muhammad Ali Jamshed

机构 * School of Electrical Engineering, Stanford University(斯坦福大学电气工程学院) College of Science and Engineering, University of Glasgow(格拉斯哥大学科学与工程学院)

专题命中 隐私与版权 :trustworthy(abstract);分类 cs.LG

Comments Submitted as a chapter to the book Ambient Intelligence for 6G

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.19876 2025-08-28 cs.SD cs.DL 50%

The IRMA Dataset: A Structured Audio-MIDI Corpus for Iranian Classical Music

Sepideh Shafiei, Shapour Hakam

机构 * Cu Test Inc.(Cu Test公司)

专题命中 隐私与版权 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

7. 安全评测 11 篇

2508.12733 2025-08-28 cs.CL cs.AI 84%

LinguaSafe: A Comprehensive Multilingual Safety Benchmark for Large Language Models

Zhiyuan Ning, Tianle Gu, Jiaxin Song, Shixin Hong, Lingyu Li, Huacan Liu, Jie Li, Yixu Wang, Meng Lingyu, Yan Teng, Yingchun Wang

机构 * Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)

专题命中 安全评测 :safety(title,abstract);alignment(abstract);分类 cs.CL、cs.AI

Comments 7pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.02531 2025-08-28 cs.CY cs.CL 81%

Towards New Benchmark for AI Alignment & Sentiment Analysis in Socially Important Issues: A Comparative Study of Human and LLMs in the Context of AGI

Ljubisa Bojic, Dylan Seychell, Milan Cabarkapa

专题命中 安全评测 :alignment(title,abstract);分类 cs.CL、cs.CY

Comments 34 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.19487 2025-08-28 cs.LG cs.AI 62%

Data-Efficient Symbolic Regression via Foundation Model Distillation

Wangyang Ying, Jinghan Zhang, Haoyue Bai, Nanxu Gong, Xinyuan Wang, Kunpeng Liu, Chandan K. Reddy, Yanjie Fu

机构 * Institute for Clarity in Documentation(文档清晰研究所) Inria Paris-Rocquencourt(巴黎-罗克琴特研究所) Rajiv Gandhi University(拉吉夫·甘地大学) Tsinghua University(清华大学) Palmer Research Laboratories(帕尔默研究实验室) Arizona State University(亚利桑那州立大学) Clemson University(克莱姆森大学)

专题命中 安全评测 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.19271 2025-08-28 cs.CL cs.AI 62%

Rethinking Reasoning in LLMs: Neuro-Symbolic Local RetoMaton Beyond ICL and CoT

Rushitha Santhoshi Mamidala, Anshuman Chhabra, Ankur Mali

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.19980 2025-08-28 cs.LG 57%

Evaluating Language Model Reasoning about Confidential Information

Dylan Sam, Alexander Robey, Andy Zou, Matt Fredrikson, J. Zico Kolter

专题命中 安全评测 :safety(abstract);分类 cs.LG

Comments 20 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.19882 2025-08-28 cs.SE cs.AI 57%

Generative AI for Testing of Autonomous Driving Systems: A Survey

Qunying Song, He Ye, Mark Harman, Federica Sarro

机构 * University College London(伦敦大学学院)

专题命中 安全评测 :safety(abstract);分类 cs.AI

Comments 67 pages, 6 figures, 29 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.19641 2025-08-28 cs.CR cs.AI 57%

Intellectual Property in Graph-Based Machine Learning as a Service: Attacks and Defenses

Lincan Li, Bolin Shen, Chenxi Zhao, Yuxiang Sun, Kaixiang Zhao, Shirui Pan, Yushun Dong

机构 * Department of Computer Science, Florida State University(佛罗里达州立大学计算机科学系) Northeastern University(东北大学) University of Wisconsin-Madison(威斯康星大学麦迪逊分校) University of Notre Dame(诺丁汉大学) School of Information and Communication Technology, Griffith University(格里菲斯大学信息与通信技术学院)

专题命中 安全评测 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.09242 2025-08-28 cs.AI 57%

From large language models to multimodal AI: A scoping review on the potential of generative AI in medicine

Lukas Buess, Matthias Keicher, Nassir Navab, Andreas Maier, Soroosh Tayebi Arasteh

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

Journal ref Biomed. Eng. Lett. 15 (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.19773 2025-08-28 cs.CV 50%

The Return of Structural Handwritten Mathematical Expression Recognition

Jakob Seitz, Tobias Lengfeld, Radu Timofte

机构 * Computer Vision Lab, CAIDAS \& IFI, University of W\"urzburg, W\"urzburg, Germany

专题命中 安全评测 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.19639 2025-08-28 cs.MM 50%

FakeSV-VLM: Taming VLM for Detecting Fake Short-Video News via Progressive Mixture-Of-Experts Adapter

Junxi Wang, Yaxiong Wang, Lechao Cheng, Zhun Zhong

专题命中 安全评测 :alignment(abstract)

Comments EMNLP2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21696 2025-08-28 eess.SP 50%

Edge Agentic AI Framework for Autonomous Network Optimisation in O-RAN

Abdelaziz Salama, Zeinab Nezami, Mohammed M. H. Qazzaz, Maryam Hafeez, Syed Ali Raza Zaidi

专题命中 安全评测 :safety(abstract)

Journal ref IEEE International Symposium on Personal, Indoor and Mobile Radio Communications (PIMRC), 2025

详情

展开后加载摘要…

URL PDF HTML 收藏

8. AI治理与伦理 3 篇

2508.19269 2025-08-28 cs.CY cs.AI cs.CL 67%

Should LLMs be WEIRD? Exploring WEIRDness and Human Rights in Large Language Models

Ke Zhou, Marios Constantinides, Daniele Quercia

机构 * Nokia Bell Labs(诺基亚贝尔实验室)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI、cs.CY

Comments This paper has been accepted in AIES 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20015 2025-08-28 cs.LG cs.AI 62%

Decomposing Behavioral Phase Transitions in LLMs: Order Parameters for Emergent Misalignment

Julian Arnold, Niels Lörch

机构 * Department of Physics University of Basel(物理系 巴塞尔大学)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.LG

Comments 11+25 pages, 4+11 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.18765 2025-08-28 cs.LG 57%

Governance-as-a-Service: A Multi-Agent Framework for AI System Compliance and Policy Enforcement

Suyash Gaurav, Jukka Heikkonen, Jatin Chaudhary

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏

9. 其他安全 8 篇

2508.19574 2025-08-28 cs.CV cs.AI 74%

Multimodal Prototype Alignment for Semi-supervised Pathology Image Segmentation

Mingxi Fu, Fanglei Fu, Xitong Ling, Huaitian Yuan, Tian Guan, Yonghong He, Lianghui Zhu

机构 * Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院)

专题命中 其他安全 :alignment(title);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20018 2025-08-28 cs.AI cs.CL cs.CV cs.MA 62%

SWIRL: A Staged Workflow for Interleaved Reinforcement Learning in Mobile GUI Control

Quanfeng Lu, Zhantao Ma, Shuai Zhong, Jin Wang, Dahai Yu, Michael K. Ng, Ping Luo

机构 * The University of Hong Kong(香港大学) Hong Kong Baptist University(香港 Baptist 大学) TCL Corporate Research (Hong Kong) Co., Ltd(TCL 香港研究院)

专题命中 其他安全 :safety(abstract);分类 cs.CL、cs.AI

Comments 28 pages, 12 figures

详情

展开后加载摘要…

URL PDF HTML 收藏