arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-11-05 至 2025-11-05 共收录 32 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 偏好对齐 1 篇

2510.06915 2025-11-05 cs.CL cs.AI 62%

LongRM: Revealing and Unlocking the Context Boundary of Reward Modeling

Zecheng Tang, Baibei Ji, Quantong Qiu, Haitian Wang, Xiaobo Liang, Juntao Li, Min Zhang

机构 * Soochow University(苏州大学) LCM Laboratory(LCM实验室)

专题命中 偏好对齐 :safety(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 安全训练 2 篇

2511.02371 2025-11-05 cs.LG 79%

LUMA-RAG: Lifelong Multimodal Agents with Provably Stable Streaming Alignment

Rohan Wandre, Yash Gajewar, Namrata Patel, Vivek Dhalkari

机构 * Dept. of Computer Engineering(计算机工程系) SIES Graduate School of Technology(SIES技术研究生学院) Bharatiya Vidya Bhavan's Sardar Patel Institute of Technology(巴哈里亚·维达·巴万学院萨达尔·帕特尔技术学院)

专题命中 安全训练 :alignment(title,abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.09024 2025-11-05 cs.CV cs.LG 57%

DIsoN: Decentralized Isolation Networks for Out-of-Distribution Detection in Medical Imaging

Felix Wagner, Pramit Saha, Harry Anthony, J. Alison Noble, Konstantinos Kamnitsas

机构 * Department of Engineering Science, University of Oxford(工程科学系,牛津大学)

专题命中 安全训练 :safety(abstract);分类 cs.LG

Comments Accepted at NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 越狱攻击 2 篇

2511.00689 2025-11-05 cs.CL 85%

Do Methods to Jailbreak and Defend LLMs Generalize Across Languages?

Berk Atil, Rebecca J. Passonneau, Fred Morstatter

机构 * Penn State University(宾夕法尼亚州立大学) Information Sciences Institute, USC(信息科学研究所)

专题命中 越狱攻击 :jailbreak(title,abstract);alignment(abstract);safety(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.02823 2025-11-05 cs.AI 57%

Optimizing AI Agent Attacks With Synthetic Data

Chloe Loughridge, Paul Colognese, Avery Griffin, Tyler Tracy, Jon Kutasov, Joe Benton

专题命中 越狱攻击 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 幻觉与事实性 1 篇

2511.01902 2025-11-05 cs.CY cs.AI 62%

Before the Clinic: Transparent and Operable Design Principles for Healthcare AI

Alexander Bakumenko, Aaron J. Masino, Janine Hoelscher

机构 * Clemson University(克莱姆森大学)

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 隐私与版权 2 篇

2510.26025 2025-11-05 cs.LG 74%

Exploring Human-AI Conceptual Alignment through the Prism of Chess

Semyon Lomasov, Judah Goldfeder, Mehmet Hamza Erol, Matthew So, Yao Yan, Addison Howard, Nathan Kutz, Ravid Shwartz Ziv

机构 * Stanford University(斯坦福大学) Columbia University(哥伦比亚大学) Kaggle University of Washington(华盛顿大学) NYU(纽约大学)

专题命中 隐私与版权 :alignment(title);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.02055 2025-11-05 cs.CR 50%

Private Map-Secure Reduce: Infrastructure for Efficient AI Data Markets

Sameer Wagh, Kenneth Stibler, Shubham Gupta, Lacey Strahm, Irina Bejan, Jiahao Chen, Dave Buckley, Ruchi Bhatia, Jack Bandy, Aayush Agarwal, Andrew Trask

专题命中 隐私与版权 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

6. 安全评测 8 篇

2511.02602 2025-11-05 quant-ph cs.AI 83%

Trustworthy Quantum Machine Learning: A Roadmap for Reliability, Robustness, and Security in the NISQ Era

Ferhat Ozgur Catak, Jungwon Seo, Umit Cali

机构 * Department of Electrical Engineering and Computer Science, University of Stavanger, Norway(斯瓦尔巴大学电气工程与计算机科学系) School of Physics, Engineering and Technology, University of York, UK(约克大学物理、工程与技术学院)

专题命中 安全评测 :trustworthy(title,abstract);safety(abstract);分类 cs.AI

Comments 22 Pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00606 2025-11-05 cs.CL 74%

SpecDiff-2: Scaling Diffusion Drafter Alignment For Faster Speculative Decoding

Jameson Sandler, Jacob K. Christopher, Thomas Hartvigsen, Ferdinando Fioretto

专题命中 安全评测 :alignment(title);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.02148 2025-11-05 cs.LG 74%

CFL: On the Use of Characteristic Function Loss for Domain Alignment in Machine Learning

Abdullah Almansour, Ozan Tonguz

机构 * Carnegie Mellon University(卡内基梅隆大学)

专题命中 安全评测 :alignment(title);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.01805 2025-11-05 cs.CL cs.AI 73%

Accumulating Context Changes the Beliefs of Language Models

Jiayi Geng, Howard Chen, Ryan Liu, Manoel Horta Ribeiro, Robb Willer, Graham Neubig, Thomas L. Griffiths

专题命中 安全评测 :alignment(abstract);safety(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.27195 2025-11-05 cs.CV cs.CL cs.SI 70%

Can MLLMs Read the Room? A Multimodal Benchmark for Verifying Truthfulness in Multi-Party Social Interactions

Caixin Kang, Yifei Huang, Liangyang Ouyang, Mingfang Zhang, Yoichi Sato

机构 * The University of Tokyo(东京大学)

专题命中 安全评测 :alignment(abstract);trustworthy(abstract);分类 cs.CL

Comments ICCV2025 Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.24414 2025-11-05 cs.CV 67%

A Quantitative Evaluation Framework for Explainable AI in Semantic Segmentation

Reem Hammoud, Abdul Karim Gizzini, Ali J. Ghandour

机构 * American University of Beirut(美国贝鲁特美国大学) SogetiLabs Research and Innovation(SogetiLabs研究与创新) National Center for Remote Sensing(远程 sensing 国家中心)

专题命中 安全评测 :safety(abstract);trustworthy(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.02495 2025-11-05 cs.CV cs.CL 57%

DetectiumFire: A Comprehensive Multi-modal Dataset Bridging Vision and Language for Fire Understanding

Zixuan Liu, Siavash H. Khajavi, Guangkai Jiang

机构 * Department of Computer Science(计算机科学系) Tulane University(Tulane 大学) Department of Industrial Engineering and Management(工业工程与管理系) Aalto University(Aalto 大学)

专题命中 安全评测 :safety(abstract);分类 cs.CL

Comments Advances in Neural Information Processing Systems 2025 (NeurIPS 2025), Poster, https://neurips.cc/virtual/2025/loc/san-diego/poster/121400

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.02223 2025-11-05 physics.med-ph 50%

Quantitative Risk Assessment in Radiation Oncology via LLM-Powered Root Cause Analysis of Incident Reports

Yuntao Wang, Siamak P. Najad-Davarani, Elizabeth Bossart, Matthew T. Studenski, Mariluz De Ornelas, Yunze Yang

专题命中 安全评测 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

7. AI治理与伦理 3 篇

2409.09586 2025-11-05 cs.HC cs.AI cs.CL 81%

ValueCompass: A Framework for Measuring Contextual Value Alignment Between Human and LLMs

Hua Shen, Tiffany Knearem, Reshmi Ghosh, Yu-Ju Yang, Nicholas Clark, Tanushree Mitra, Yun Huang

机构 * NYU Shanghai(纽约大学上海校区) New York University(纽约大学) University of Washington(华盛顿大学) MBZUAI Microsoft(微软) UIUC(伊利诺伊大学香槟分校)

专题命中 AI治理与伦理 :alignment(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15975 2025-11-05 cs.CR q-bio.BM 80%

Generative AI for Biosciences: Emerging Threats and Roadmap to Biosecurity

Zaixi Zhang, Souradip Chakraborty, Amrit Singh Bedi, Emilin Mathew, Varsha Saravanan, Le Cong, Alvaro Velasquez, Sheng Lin-Gibson, Megan Blewett, Dan Hendrycs, Alex John London, Ellen Zhong, Ben Raphael, Adji Bousso Dieng, Jian Ma, Eric Xing, Russ Altman, George Church, Mengdi Wang

专题命中 AI治理与伦理 :alignment(abstract);safety(abstract);jailbreak(abstract);AI safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.25863 2025-11-05 cs.CR cs.AI 57%

AAGATE: A NIST AI RMF-Aligned Governance Platform for Agentic AI

Ken Huang, Kyriakos Rock Lambros, Jerry Huang, Yasir Mehmood, Hammad Atta, Joshua Beck, Vineeth Sai Narajala, Muhammad Zeeshan Baig, Muhammad Aziz Ul Haq, Nadeem Shahzad, Bhavya Gupta

机构 * RockCyber Kleiner Perkins Qorvex Consulting & Roshan Consulting SAS Institute OWASP Wentworth Institute of Higher Education & Machine Learning Professional Skylink Antenna Roshan Consulting & Robotic Process Automation Stanford University

专题命中 AI治理与伦理 :red teaming(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏

8. 其他安全 13 篇

2509.04104 2025-11-05 cs.CL cs.HC 79%

Towards Stable and Personalised Profiles for Lexical Alignment in Spoken Human-Agent Dialogue

Keara Schaaij, Roel Boumans, Tibor Bosse, Iris Hendrickx

机构 * Centre for Language Studies, Centre for Language and Speech Technology, Radboud University,Nijmegen, The Netherlands(语言研究所以及语言与语音技术中心,拉德堡德大学,尼姆egen,荷兰) Behavioural Science Institute, Radboud University, Nijmegen, The Netherlands(行为科学研究所,拉德堡德大学,尼姆egen,荷兰)

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL

Comments This preprint has not undergone peer review or any post-submission improvements or corrections. The Version of Record of this contribution is published in TSD 2025. Lecture Notes in Computer Science, vol 16029

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.09485 2025-11-05 cs.RO cs.AI cs.GR 79%

Adv-BMT: Bidirectional Motion Transformer for Safety-Critical Traffic Scenario Generation

Yuxin Liu, Zhenghao Peng, Xuanhao Cui, Bolei Zhou

机构 * University of California, Los Angeles(加州大学洛杉矶分校)

专题命中 其他安全 :safety(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.18835 2025-11-05 cs.SE 71%

AUCAD: Automated Construction of Alignment Dataset from Log-Related Issues for Enhancing LLM-based Log Generation

Hao Zhang, Dongjun Yu, Lei Zhang, Guoping Rong, Yongda Yu, Haifeng Shen, He Zhang, Dong Shao, Hongyu Kuang

专题命中 其他安全 :alignment(title)

Comments In the 16th International Conference on Internetware 2025. 13 pages

Journal ref Proceedings of the 16th International Conference on Internetware (2025) 413-425

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.01894 2025-11-05 cs.GR cs.AI cs.LG 62%

LGCC: Enhancing Flow Matching Based Text-Guided Image Editing with Local Gaussian Coupling and Context Consistency

Fangbing Liu, Pengfei Duan, Wen Li, Yi He

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.05862 2025-11-05 cs.CL cs.AI 62%

Revisiting Long-context Modeling from Context Denoising Perspective

Zecheng Tang, Baibei Ji, Juntao Li, Lijun Wu, Haijia Gui, Min Zhang

机构 * Soochow University(苏州大学) LCM Laboratory(长文实验室) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.15090 2025-11-05 cs.CL cs.AI 62%

ExpertLens: Activation steering features are highly interpretable

Masha Fedzechkina, Eleonora Gualdoni, Sinead Williamson, Katherine Metcalf, Skyler Seto, Barry-John Theobald

机构 * Apple(苹果公司)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.02690 2025-11-05 cs.LG 57%

Curriculum Design for Trajectory-Constrained Agent: Compressing Chain-of-Thought Tokens in LLMs

Georgios Tzannetos, Parameswaran Kamalaruban, Adish Singla

机构 * MPI-SWS(马克斯·普朗克所际研究所)

专题命中 其他安全 :safety(abstract);分类 cs.LG

Comments NeurIPS'25 paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.02606 2025-11-05 cs.AI cs.HC 57%

A Multi-Agent Psychological Simulation System for Human Behavior Modeling

Xiangen Hu, Jiarui Tong, Sheng Xu

专题命中 其他安全 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21971 2025-11-05 cs.LG 57%

GRAM-DTI: adaptive multimodal representation learning for drug target interaction prediction

Feng Jiang, Amina Mollaysa, Hehuan Ma, Tommaso Mansi, Junzhou Huang, Mangal Prakash, Rui Liao

机构 * University of Texas at Arlington(德克萨斯理工大学) Johnson & Johnson Innovative Medicine(强生创新医学)

专题命中 其他安全 :alignment(abstract);分类 cs.LG

Journal ref NeurIPS 2025 2nd Workshop on Multi-modal Foundation Models and Large Language Models for Life Sciences

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.06910 2025-11-05 cs.CL 57%

Identifying Aspects in Peer Reviews

Sheng Lu, Ilia Kuznetsov, Iryna Gurevych

机构 * Ubiquitous Knowledge Processing Lab (UKP Lab)(通用知识处理实验室) Department of Computer Science(计算机科学系) Hessian Center for AI (hessian.AI)(黑森人工智能中心)

专题命中 其他安全 :alignment(abstract);分类 cs.CL

Comments EMNLP 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.02711 2025-11-05 cs.DB cs.IR 50%

Relational Deep Dive: Error-Aware Queries Over Unstructured Data

Daren Chao, Kaiwen Chen, Naiqing Guan, Nick Koudas

专题命中 其他安全 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏