arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-11-10 至 2025-11-10 共收录 36 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 偏好对齐 5 篇

2508.18312 2025-11-10 cs.LG cs.AI 84%

What Matters in Data for DPO?

Yu Pan, Zhongze Cai, Guanting Chen, Huaiyang Zhong, Chonghuan Wang

机构 * University of Sydney(悉尼大学) Imperial College London(伦敦帝国理工学院) University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校) Virginia Tech(弗吉尼亚理工大学) University of Texas at Dallas(德克萨斯大学达拉斯分校)

专题命中 偏好对齐 :DPO(title,abstract);alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.05359 2025-11-10 cs.CR cs.CL cs.CY 81%

ConVerse: Benchmarking Contextual Safety in Agent-to-Agent Conversations

Amr Gomaa, Ahmed Salem, Sahar Abdelnabi

机构 * German Research Center for Artificial Intelligence (DFKI)(德国人工智能研究中心(DFKI)) Microsoft(微软) ELLIS Institute Tübingen and MPI for Intelligent Systems(图宾根ELLIS研究所和智能系统研究所) Tübingen AI Center(图宾根人工智能中心)

专题命中 偏好对齐 :safety(title,abstract);分类 cs.CL、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.04880 2025-11-10 cs.AI 79%

DMA: Online RAG Alignment with Human Feedback

Yu Bai, Yukai Miao, Dawei Wang, Li Chen, Fei Long, Rundi Zhai, Dan Li, Yanyu Ren, Tianfeng Liu, Hongtao Xie, Ce Yang, Xuhui Cai

机构 * Zhongguancun Laboratory(中关村实验室) Tsinghua University(清华大学) China Mobile Communications Group Co., Ltd.(中国移动通信集团有限公司) Beijing University of Posts and Telecommunications(北京邮电大学)

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.04845 2025-11-10 cs.LG 57%

Investigating U.S. Consumer Demand for Food Products with Innovative Transportation Certificates Based on Stated Preferences and Machine Learning Approaches

Jingchen Bi, Rodrigo Mesa-Arango

专题命中 偏好对齐 :safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.03740 2025-11-10 cs.CL 57%

LEME: Open Large Language Models for Ophthalmology with Advanced Reasoning and Clinical Validation

Hyunjae Kim, Xuguang Ai, Sahana Srinivasan, Aidan Gilson, Maxwell B. Singer, Krithi Pushpanathan, Qianqian Xie, Jungwoo Park, Serina Applebaum, Gabriel Dawei Yang, Minjie Zou, David Ziyou Chen, Ke Zou, Soshian Sarrafpour, Ji Liu, Yu Yin, Jimin Huang, Quang Ngoc Nguyen, Erping Long, Peixing Wan, Dianbo Liu, Richard Hintz, W. Jim Zheng, Sophia Y. Wang, Lucila Ohno-Machado, Hua Xu, Ron A. Adelman, Luciano V. Del Priore, Yih-Chung Tham, Qingyu Chen

专题命中 偏好对齐 :safety(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 安全训练 3 篇

2410.14865 2025-11-10 cs.AI cs.FL cs.RO 79%

Joint Verification and Refinement of Language Models for Safety-Constrained Planning

Yunhao Yang, Neel P. Bhatt, William Ward, Zichao Hu, Joydeep Biswas, Ufuk Topcu

机构 * University of Texas at Austin(德克萨斯大学奥斯汀分校)

专题命中 安全训练 :safety(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.05018 2025-11-10 cs.CL cs.AI cs.LG 75%

Pluralistic Behavior Suite: Stress-Testing Multi-Turn Adherence to Custom Behavioral Policies

Prasoon Varshney, Makesh Narsimhan Sreedhar, Liwei Jiang, Traian Rebedea, Christopher Parisien

机构 * NVIDIA

专题命中 安全训练 :alignment(abstract);safety(abstract);分类 cs.CL、cs.AI、cs.LG

Comments Accepted at the Multi-Turn Interactions workshop at the 39th Conference on Neural Information Processing Systems (NeurIPS 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.05286 2025-11-10 cs.CL 57%

Reflective Personalization Optimization: A Post-hoc Rewriting Framework for Black-Box Large Language Models

Teqi Hao, Xioayu Tan, Shaojie Shi, Yinghui Xu, Xihe Qiu

机构 * School of Electronic and Electrical Engineering, Shanghai University of Engineering Science(上海工程技术大学电子与电气工程学院) Tencent Youtu Lab(腾讯优图实验室) Artificial Intelligence Innovation and Incubation Institute, Fudan University(复旦大学人工智能创新与孵化院)

专题命中 安全训练 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 越狱攻击 3 篇

2504.21700 2025-11-10 cs.CR cs.AI cs.LG 81%

XBreaking: Understanding how LLMs security alignment can be broken

Marco Arazzi, Vignesh Kumar Kembu, Antonino Nocera, Vinod P

机构 * Department of Electrical, Computer Biomedical Engineering, University of Pavia, Italy\ . Department of Computer Applications,\ University of Science \& Technology, India\ .

专题命中 越狱攻击 :alignment(title,abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.18469 2025-11-10 cs.CL cs.LG 79%

Iterative Self-Tuning LLMs for Enhanced Jailbreaking Capabilities

Chung-En Sun, Xiaodong Liu, Weiwei Yang, Tsui-Wei Weng, Hao Cheng, Aidan San, Michel Galley, Jianfeng Gao

专题命中 越狱攻击 :alignment(abstract);safety(abstract);jailbreak(abstract);分类 cs.CL、cs.LG

Comments Accepted to NAACL 2025 Main (Oral)

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.03299 2025-11-10 cs.LG cs.CL cs.CV 62%

GUARD: Role-playing to Generate Natural-language Jailbreakings to Test Guideline Adherence of Large Language Models

Haibo Jin, Ruoxi Chen, Peiyan Zhang, Andy Zhou, Haohan Wang

机构 * School of Information Sciences University of Illinois at Urbana-Champaign(信息科学学院伊利诺伊大学厄巴纳-香槟分校) Independent Researcher, Starc Institute(Starc研究所独立研究者) Computer Science and Engineering HKUST(HKUST计算机科学与工程学院) Computer Science Lapis Labs University of Illinois Urbana-Champaign(计算机科学Lapis Labs伊利诺伊大学厄巴纳-香槟分校) School of Information Sciences University of Illinois Urbana-Champaign(信息科学学院伊利诺伊大学厄巴纳-香槟分校)

专题命中 越狱攻击 :safety(abstract);分类 cs.CL、cs.LG

Comments 28 papges

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 幻觉与事实性 2 篇

2509.09360 2025-11-10 cs.CL 57%

MetaRAG: Metamorphic Testing for Hallucination Detection in RAG Systems

Channdeth Sok, David Luz, Yacine Haddam

机构 * Forvia Paris Tech Center, GIT, Immeuble Lumière, 40 avenue des Terroirs de France, 75012 Paris, France(巴黎Forvia技术中心,GIT,Lumière大厦,法国巴黎75012)

专题命中 幻觉与事实性 :trustworthy(abstract);分类 cs.CL

Comments Identity-Aware AI workshop at 28th European Conference on Artificial Intelligence, October 25, 2025, Bologna, Italy

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.06645 2025-11-10 q-bio.NC 50%

Quantifying Uncertainty in Error Consistency: Towards Reliable Behavioral Comparison of Classifiers

Thomas Klein, Sascha Meyen, Wieland Brendel, Felix A. Wichmann, Kristof Meding

专题命中 幻觉与事实性 :alignment(abstract)

Comments 25 pages, 14 figures

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 安全评测 14 篇

2501.03265 2025-11-10 cs.LG cs.AI 73%

Cognitive Edge Computing: A Comprehensive Survey on Optimizing Large Models and AI Agents for Pervasive Deployment

Xubin Wang, Qing Li, Weijia Jia

专题命中 安全评测 :alignment(abstract);safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.04886 2025-11-10 cs.CV cs.AI 70%

Beta Distribution Learning for Reliable Roadway Crash Risk Assessment

Ahmad Elallaf, Nathan Jacobs, Xinyue Ye, Mei Chen, Gongbo Liang

机构 * Texas A&M University-San Antonio(德克萨斯A&M大学-圣安东尼奥分校) Washington University in St. Louis(华盛顿大学圣路易斯分校) University of Alabama(阿拉巴马大学) University of Kentucky(肯塔基大学)

专题命中 安全评测 :safety(abstract);trustworthy(abstract);分类 cs.AI

Comments Accepted to AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15144 2025-11-10 cs.AI cs.CL cs.CY 67%

HugAgent: Benchmarking LLMs for Simulation of Individualized Human Reasoning

Chance Jiajie Li, Zhenze Mo, Yuhan Tang, Ao Qu, Jiayi Wu, Kaiya Ivy Zhao, Yulu Gan, Jie Fan, Jiangbo Yu, Hang Jiang, Paul Pu Liang, Jinhua Zhao, Luis Alberto Alonso Pastor, Kent Larson

机构 * MIT Media Lab(MIT媒体实验室) MIT EECS(MIT电子工程与计算机科学系) MIT BCS(MIT生物科学系) MIT IDSS(MIT国际设计系统研究所) MIT CEE(MIT土木与环境工程系) MIT DUSP(MIT设计学教授职位) MIT Architecture(MIT建筑系) Northeastern University(东北大学) Brown University(布朗大学) McGill University(麦吉尔大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI、cs.CY

Comments To appear in NeurIPS 2025 Workshop on Bridging Language, Agent, and World Models (LAW)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.17797 2025-11-10 cs.CL cs.AI 62%

Enterprise Deep Research: Steerable Multi-Agent Deep Research for Enterprise Analytics

Akshara Prabhakar, Roshan Ram, Zixiang Chen, Silvio Savarese, Frank Wang, Caiming Xiong, Huan Wang, Weiran Yao

机构 * Salesforce AI Research(Salesforce AI研究院)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

Comments Technical report; 13 pages plus references and appendices

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.04956 2025-11-10 cs.AI cs.CL 62%

ORCHID: Orchestrated Retrieval-Augmented Classification with Human-in-the-Loop Intelligent Decision-Making for High-Risk Property

Maria Mahbub, Vanessa Lama, Sanjay Das, Brian Starks, Christopher Polchek, Saffell Silvers, Lauren Deck, Prasanna Balaprakash, Tirthankar Ghosal

机构 * Oak Ridge National Laboratory(橡树岭国家实验室) Pacific Northwest National Laboratory(太平洋西北国家实验室)

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.04703 2025-11-10 cs.CL cs.AI 62%

Measuring what Matters: Construct Validity in Large Language Model Benchmarks

Andrew M. Bean, Ryan Othniel Kearns, Angelika Romanou, Franziska Sofia Hafner, Harry Mayne, Jan Batzner, Negar Foroutan, Chris Schmitz, Karolina Korgul, Hunar Batra, Oishi Deb, Emma Beharry, Cornelius Emde, Thomas Foster, Anna Gausen, María Grandury, Simeng Han, Valentin Hofmann, Lujain Ibrahim, Hazel Kim, Hannah Rose Kirk, Fangru Lin, Gabrielle Kaili-May Liu, Lennart Luettgau, Jabez Magomere, Jonathan Rystrøm, Anna Sotnikova, Yushi Yang, Yilun Zhao, Adel Bibi, Antoine Bosselut, Ronald Clark, Arman Cohan, Jakob Foerster, Yarin Gal, Scott A. Hale, Inioluwa Deborah Raji, Christopher Summerfield, Philip H. S. Torr, Cozmin Ududec, Luc Rocher, Adam Mahdi

机构 * University of Oxford(牛津大学) EPFL(苏黎世联邦理工学院) Weizenbaum Institute Berlin(柏林Weizenbaum研究所) Technical University Munich(慕尼黑技术大学) Centre for Digital Governance, Hertie School(赫尔姆霍兹学院数字治理中心) Stanford University(斯坦福大学) UK AI Security Institute(英国人工智能安全研究所) SomosNLP Universdad Politécnica de Madrid(马德里理工大学) Yale University(耶鲁大学) Allen Institute for AI(人工智能研究所) University of Washington(华盛顿大学) Meedan UC Berkeley(加州大学伯克利分校)

专题命中 安全评测 :safety(abstract);分类 cs.CL、cs.AI

Comments 39th Conference on Neural Information Processing Systems (NeurIPS 2025) Track on Datasets and Benchmarks

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.04247 2025-11-10 cs.MM cs.AI cs.IR 57%

On the Brittleness of CLIP Text Encoders

Allie Tran, Luca Rossetto

机构 * Dublin City University(都柏林城市大学)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

Comments Accepted for publication at MMM'26. Analysis code can be found here: https://github.com/allie-tran/clip-brittleness

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.05299 2025-11-10 cs.CV cs.AI 57%

LiveStar: Live Streaming Assistant for Real-World Online Video Understanding

Zhenyu Yang, Kairui Zhang, Yuhang Hu, Bing Wang, Shengsheng Qian, Bin Wen, Fan Yang, Tingting Gao, Weiming Dong, Changsheng Xu

机构 * Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) University of Chinese Academy of Sciences(中国科学院大学) ShanghaiTech University(上海科技大学) Kuaishou Technology(快手科技) Peng Cheng Laboratory(鹏城实验室)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

Comments NeurIPS 2025 Accepted

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.05269 2025-11-10 cs.MA cs.AI 57%

TAMAS: Benchmarking Adversarial Risks in Multi-Agent LLM Systems

Ishan Kavathekar, Hemang Jain, Ameya Rathod, Ponnurangam Kumaraguru, Tanuja Ganu

机构 * International Institute of Information Technology, Hyderabad(国际信息科技研究所,海得拉巴) Microsoft Research, India(微软研究院,印度)

专题命中 安全评测 :safety(abstract);分类 cs.AI

Comments Accepted at ICML 2025 MAS Workshop. This version includes additional experiments and analysis

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.05000 2025-11-10 cs.IR cs.AI 57%

Query Generation Pipeline with Enhanced Answerability Assessment for Financial Information Retrieval

Hyunkyu Kim, Yeeun Yoo, Youngjun Kwak

机构 * Kakaobank Seongnam-si Republic of Korea(韩国首尔市韩国 Kakao银行)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

Comments Accepted(Oral) by ICAIF 2025. Hyunkyu Kim and Yeeun Yoo contributed equally to this work

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.04690 2025-11-10 cs.MM cs.CL 57%

Automatización de Informes Geotécnicos para Macizos Rocosos con IA

Christofer Valencia, Alexis Llumigusín, Silvia Alvarez, Abrahan Arias, Christian Mejia-Escobar

专题命中 安全评测 :safety(abstract);分类 cs.CL

Comments 17 pages, in Spanish language

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21069 2025-11-10 cs.CV cs.AI cs.HC 57%

GAITEX: Human motion dataset of impaired gait and rehabilitation exercises using inertial and optical sensors

Andreas Spilz, Heiko Oppel, Jochen Werner, Kathrin Stucke-Straub, Felix Capanni, Michael Munz

机构 * AI for Sensor Data Analytics Research Group(人工智能传感器数据解析研究组) Ulm University of Applied Sciences(乌尔姆应用科学大学) Biomechatronic Research Group(生物机械研究组) Institute of Computer Science(计算机科学研究所)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.11137 2025-11-10 cs.CL 57%

Scalable Medication Extraction and Discontinuation Identification from Electronic Health Records Using Large Language Models

Chong Shao, Douglas Snyder, Chiran Li, Bowen Gu, Kerry Ngan, Chun-Ting Yang, Jiageng Wu, Richard Wyss, Kueiyu Joshua Lin, Jie Yang

专题命中 安全评测 :safety(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.17086 2025-11-10 cs.CL 57%

Mind the Blind Spots: A Focus-Level Evaluation Framework for LLM Reviews

Hyungyu Shin, Jingyu Tang, Yoonjoo Lee, Nayoung Kim, Hyunseung Lim, Ji Yong Cho, Hwajung Hong, Moontae Lee, Juho Kim

机构 * KAIST(韩国科学技术院) Huazhong University of Science and Technology(华中科技大学) LG AI Research(LG人工智能研究) University of Illinois Chicago(伊利诺伊大学香槟分校)

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL

Comments EMNLP 2025 Oral

详情

展开后加载摘要…

URL PDF HTML 收藏

6. AI治理与伦理 1 篇

2509.23994 2025-11-10 cs.CL cs.AI 73%

Policy-as-Prompt: Turning AI Governance Rules into Guardrails for AI Agents

Gauri Kholkar, Ratinder Ahuja

机构 * Pure Storage

专题命中 AI治理与伦理 :safety(abstract);AI safety(abstract);分类 cs.CL、cs.AI

Comments Accepted at 3rd Regulatable ML Workshop at NEURIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏

7. 其他安全 8 篇

2410.02615 2025-11-10 cs.LG 79%

ExGra-Med: Extended Context Graph Alignment for Medical Vision-Language Models

Duy M. H. Nguyen, Nghiem T. Diep, Trung Q. Nguyen, Hoang-Bao Le, Tai Nguyen, Tien Nguyen, TrungTin Nguyen, Nhat Ho, Pengtao Xie, Roger Wattenhofer, James Zou, Daniel Sonntag, Mathias Niepert

机构 * German Research Centre for Artificial Intelligence (DFKI)(德国人工智能研究中心) Max Planck Research School for Intelligent Systems (IMPRS-IS)(马克斯·普朗克智能系统研究学校) University of Stuttgart(斯图加特大学) University Medical Center Gottingen(哥廷根大学医学中心) Max Planck Institute for Multidisciplinary Sciences(马克斯·普朗克多学科科学研究所) ARC Centre of Excellence for the Mathematical Analysis of Cellular Systems(细胞系统数学分析卓越中心) School of Mathematical Sciences, Queensland University of Technology(昆士兰科技大学数学科学学院) University of Oldenburg(奥尔登堡大学) University of Texas at Austin(德克萨斯大学奥斯汀分校) University of California San Diego(加州大学圣地亚哥分校) MBZUAI(马克斯·普朗克人工智能研究所) ETH Zurich(苏黎世联邦理工学院) Stanford University(斯坦福大学)

专题命中 其他安全 :alignment(title,abstract);分类 cs.LG

Comments Accepted at NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.05718 2025-11-10 cs.HC 71%

Do Vision-Language Models See Visualizations Like Humans? Alignment in Chart Categorization

Péter Ferenc Gyarmati, Manfred Klaffenböck, Laura Koesten, Torsten Möller

专题命中 其他安全 :alignment(title)

Comments 2 pages, 2 figures. Accepted submission to the poster track of IEEE VIS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏