arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-10-31 至 2025-10-31 共收录 33 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 偏好对齐 4 篇

2503.03710 2025-10-31 cs.CL cs.CR cs.LG 91%

Improving LLM Safety Alignment with Dual-Objective Optimization

Xuandong Zhao, Will Cai, Tianneng Shi, David Huang, Licong Lin, Song Mei, Dawn Song

机构 * University of California, Berkeley(加州大学伯克利分校)

专题命中 偏好对齐 :alignment(title,abstract);safety(title,abstract);DPO(abstract);jailbreak(abstract)

Comments ICML 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.25884 2025-10-31 cs.AI cs.CL cs.LG 67%

Approximating Human Preferences Using a Multi-Judge Learned System

Eitán Sprejer, Fernando Avalos, Augusto Bernardi, Jose Pedro Brito de Azevedo Faustino, Jacob Haimes, Narmeen Fatimah Oozeer

机构 * BAISH | UBA | Apart Research(BAISH | UBA | Apart研究) University of São Paulo(圣保罗大学) Apart Research(Apart研究) Dovetail Research | Apart Research(Dovetail研究 | Apart研究)

专题命中 偏好对齐 :RLHF(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.02820 2025-10-31 cs.AI cs.CL cs.LG 67%

AutoLibra: Agent Metric Induction from Open-Ended Human Feedback

Hao Zhu, Phil Cuvin, Xinkai Yu, Charlotte Ka Yee Yan, Jason Zhang, Diyi Yang

机构 * Stanford University(斯坦福大学) University of Toronto(多伦多大学) University of Pennsylvania(宾夕法尼亚大学)

专题命中 偏好对齐 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments https://github.com/Open-Social-World/autolibra

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.07127 2025-10-31 cs.RO cs.AI 57%

Human-assisted Robotic Policy Refinement via Action Preference Optimization

Wenke Xia, Yichu Yang, Hongtao Wu, Xiao Ma, Tao Kong, Di Hu

机构 * Gaoling School of Artificial Intelligence, Renmin University of China, Beijing(中国人民大学北京校区人工智能学院) Engineering Research Center of Next-Generation Intelligent Search(下一代智能搜索与推荐工程研究中心) Beijing Key Laboratory of Research on Large Models(北京大型模型研究重点实验室)

专题命中 偏好对齐 :alignment(abstract);分类 cs.AI

Comments Accepted By NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 安全训练 1 篇

2510.24871 2025-10-31 eess.SY cs.SY 71%

Decentralized Merging Control of Connected and Automated Vehicles to Enhance Safety and Energy Efficiency using Control Barrier Functions

Shreshta Rajakumar Deshpande, Mrdjan Jankovic

专题命中 安全训练 :safety(title)

Comments This work has been submitted to a conference for possible publication and is under review. Paper summary: 8 pages, 5 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 越狱攻击 2 篇

2510.26096 2025-10-31 cs.SD cs.CR cs.LG 83%

ALMGuard: Safety Shortcuts and Where to Find Them as Guardrails for Audio-Language Models

Weifei Jin, Yuxin Cao, Junjie Su, Minhui Xue, Jie Hao, Ke Xu, Jin Song Dong, Derui Wang

机构 * Beijing University of Posts and Telecommunications(北京邮电大学) National University of Singapore(新加坡国立大学) CSIRO’s Data61(CSIRO数据61) Responsible AI Research (RAIR) Centre, The University of Adelaide(负责任人工智能研究(RAIR)中心,阿德莱德大学) Tsinghua University(清华大学)

专题命中 越狱攻击 :safety(title,abstract);jailbreak(abstract);分类 cs.LG

Comments Accepted to NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08604 2025-10-31 cs.CL cs.AI cs.LG 75%

LatentBreak: Jailbreaking Large Language Models through Latent Space Feedback

Raffaele Mura, Giorgio Piras, Kamilė Lukošiūtė, Maura Pintor, Amin Karbasi, Battista Biggio

机构 * University of Cagliari(卡利亚里大学) Centre for AI Governance(人工智能治理中心) Foundation AI – Cisco Systems Inc.(AI基金会——思科系统公司)

专题命中 越狱攻击 :safety(abstract);jailbreak(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 提示注入 1 篇

2510.26328 2025-10-31 cs.LG 79%

Agent Skills Enable a New Class of Realistic and Trivially Simple Prompt Injections

David Schmotz, Sahar Abdelnabi, Maksym Andriushchenko

机构 * ELLIS Institute Tübingen(图宾根ELLIS研究所) MPI for Intelligent Systems Tübingen(图宾根智能系统研究所) AI Center(人工智能中心)

专题命中 提示注入 :prompt injection(title,abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 幻觉与事实性 1 篇

2510.26769 2025-10-31 cs.CV cs.LG 57%

SteerVLM: Robust Model Control through Lightweight Activation Steering for Vision Language Models

Anushka Sivakumar, Andrew Zhang, Zaber Hakim, Chris Thomas

机构 * Department of Computer Science(计算机科学系) Virginia Tech(弗吉尼亚理工学院)

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏

6. 安全评测 13 篇

2510.26024 2025-10-31 cs.CL cs.AI 81%

Rethinking Cross-lingual Alignment: Balancing Transfer and Cultural Erasure in Multilingual LLMs

HyoJung Han, Sweta Agrawal, Eleftheria Briakou

专题命中 安全评测 :alignment(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.08525 2025-10-31 cs.LG cs.AI 76%

A mathematical certification for positivity conditions in Neural Networks with applications to partial monotonicity and Trustworthy AI

Alejandro Polo-Molina, David Alfaya, Jose Portela

机构 * CDTI

专题命中 安全评测 :trustworthy(title);分类 cs.AI、cs.LG

Comments 16 pages, 4 figures

Journal ref IEEE Transactions on Neural Networks and Learning Systems, Early Access, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.14681 2025-10-31 cs.CL 74%

Massive Supervised Fine-tuning Experiments Reveal How Data, Layer, and Training Factors Shape LLM Alignment Quality

Yuto Harada, Yusuke Yamauchi, Yusuke Oda, Yohei Oseki, Yusuke Miyao, Yu Takagi

机构 * NII LLMC(日本信息处理学会大语言模型中心) The University of Tokyo(东京大学) NAIST(日本科学技术大学) Nagoya Institute of Technology(名古屋技术大学)

专题命中 安全评测 :alignment(title);分类 cs.CL

Comments Accepted to EMNLP 2025 (Main Conference). Models and evaluation results available at: https://github.com/llm-jp/massive-sft

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.25908 2025-10-31 cs.AI 70%

SciTrust 2.0: A Comprehensive Framework for Evaluating Trustworthiness of Large Language Models in Scientific Applications

Emily Herron, Junqi Yin, Feiyi Wang

机构 * Oak Ridge National Laboratory(橡树岭国家实验室)

专题命中 安全评测 :safety(abstract);trustworthy(abstract);分类 cs.AI

Comments Preprint Submitted to ACM Transactions on AI for Science (TAIS)

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.02927 2025-10-31 eess.SY cs.AI cs.LG cs.SY 62%

Multivariate Physics-Informed Convolutional Autoencoder for Anomaly Detection in Power Distribution Systems with High Penetration of DERs

Mehdi Jabbari Zideh, Sarika Khushalani Solanki

机构 * Lane Department of Computer Science and Electrical Engineering, West Virginia University(计算机科学与电气工程系,西弗吉尼亚大学)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG

Journal ref Sustainable Energy, Grids and Networks, Vol. 44, December 2025, 102022

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.26402 2025-10-31 cs.AI cs.LG 62%

Autograder+: A Multi-Faceted AI Framework for Rich Pedagogical Feedback in Programming Education

Vikrant Sahu, Gagan Raj Gupta, Raghav Borikar, Nitin Mane

机构 * Indian Institute of Technology(印度理工学院)

专题命中 安全评测 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.26037 2025-10-31 cs.CR cs.AI cs.CL 62%

SIRAJ: Diverse and Efficient Red-Teaming for LLM Agents via Distilled Structured Reasoning

Kaiwen Zhou, Ahmed Elgohary, A S M Iftekhar, Amin Saied

机构 * Microsoft Responsible AI Research(微软负责任人工智能研究) University of California(加州大学)

专题命中 安全评测 :safety(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.21497 2025-10-31 cs.CV cs.AI cs.CL cs.MA 62%

Paper2Poster: Towards Multimodal Poster Automation from Scientific Papers

Wei Pang, Kevin Qinghong Lin, Xiangru Jian, Xi He, Philip Torr

机构 * University of Waterloo(滑铁卢大学) University of Oxford(牛津大学) Vector Institute(向量研究所)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

Comments Project Page: https://github.com/Paper2Poster/Paper2Poster

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.26052 2025-10-31 cs.CV cs.AI 57%

Dynamic VLM-Guided Negative Prompting for Diffusion Models

Hoyeon Chang, Seungjin Kim, Yoonseok Choi

机构 * KAIST(韩国科学技术院)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

Comments 39th Conference on Neural Information Processing Systems (NeurIPS 2025) Workshop: The First Workshop on Generative and Protective AI for Content Creation

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.05747 2025-10-31 cs.CL 57%

SEA-LION: Southeast Asian Languages in One Network

Raymond Ng, Thanh Ngan Nguyen, Yuli Huang, Ngee Chia Tai, Wai Yi Leong, Wei Qi Leong, Xianbin Yong, Jian Gang Ngui, Yosephine Susanto, Nicholas Cheng, Hamsawardhini Rengarajan, Peerat Limkonchotiwat, Adithya Venkatadri Hulagadri, Kok Wai Teng, Yeo Yeow Tong, Bryan Siow, Wei Yi Teo, Wayne Lau, Choon Meng Tan, Brandon Ong, Zhi Hao Ong, Jann Railey Montalan, Adwin Chan, Sajeban Antonyrex, Ren Lee, Esther Choa, David Ong Tat-Wee, Bing Jie Darius Liu, William Chandra Tjhi, Erik Cambria, Leslie Teo

机构 * AI Singapore National University of Singapore(新加坡国立大学) Nanyang Technological University(南洋理工大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

Comments Accepted at IJCNLP-AACL 2025 (Main Track). We released our model at https://huggingface.co/collections/aisingapore/sea-lionv3-672589a39cdadd6a5b199581

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.04852 2025-10-31 cs.CV cs.LG 57%

CAUSAL3D: A Comprehensive Benchmark for Causal Learning from Visual Data

Disheng Liu, Yiran Qiao, Wuche Liu, Yiren Lu, Yunlai Zhou, Tuo Liang, Yu Yin, Jing Ma

机构 * Case Western Reserve University(凯斯西储大学)

专题命中 安全评测 :trustworthy(abstract);分类 cs.LG

Comments Datasets link: https://huggingface.co/datasets/LLDDSS/Causal3D_Dataset

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.06771 2025-10-31 cs.CV 50%

D-HUMOR: Dark Humor Understanding via Multimodal Open-ended Reasoning -- A Benchmark Dataset and Method

Sai Kartheek Reddy Kasu, Mohammad Zia Ur Rehman, Shahid Shafi Dar, Rishi Bharat Junghare, Dhanvin Sanjay Namboodiri, Nagendra Kumar

机构 * Indian Institute of Information Technology Dharwad, India(印度达拉瓦德信息科技学院) Indian Institute of Technology Indore, India(印度印度理工学院) Malaviya National Institute of Technology Jaipur, India(马拉维亚国家理工学院)

专题命中 安全评测 :alignment(abstract)

Comments Accepted at IEEE International Conference on Data Mining (ICDM) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.25402 2025-10-31 cs.IR cs.CE 50%

Towards Automated Quality Assurance of Patent Specifications: A Multi-Dimensional LLM Framework

Yuqian Chai, Chaochao Wang, Weilei Wang

专题命中 安全评测 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

7. AI治理与伦理 2 篇

2505.10603 2025-10-31 cs.CY cs.AI 62%

Toward a Public and Secure Generative AI: A Comparative Analysis of Open and Closed LLMs

Jorge Machado

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.26309 2025-10-31 cs.AI cs.IR 57%

GraphCompliance: Aligning Policy and Context Graphs for LLM-Based Regulatory Compliance

Jiseong Chung, Ronny Ko, Wonchul Yoo, Makoto Onizuka, Sungmok Kim, Tae-Wan Kim, Won-Yong Shin

机构 * Seoul National University(首尔国立大学) Osaka University(大阪大学) Yonsei University(延世大学)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI

Comments Under review at The Web Conference 2026 (Semantics & Knowledge track). Code will be released upon acceptance. This arXiv v1 contains no repository links to preserve double-blind review

详情

展开后加载摘要…

URL PDF HTML 收藏

8. 其他安全 9 篇

2510.26580 2025-10-31 cs.CV 78%

Dynamic Context-Aware Scene Reasoning Using Vision-Language Alignment in Zero-Shot Real-World Scenarios

Manjunath Prasad Holenarasipura Rajiv, B. M. Vidyavathi

专题命中 其他安全 :alignment(title,abstract)

Comments Preprint under review at IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI), 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.26464 2025-10-31 cs.CV 78%

Towards Fine-Grained Vision-Language Alignment for Few-Shot Anomaly Detection

Yuanting Fan, Jun Liu, Xiaochen Chen, Bin-Bin Gao, Jian Li, Yong Liu, Jinlong Peng, Chengjie Wang

机构 * Tencent Youtu Lab(腾讯优图实验室)

专题命中 其他安全 :alignment(title,abstract)

Comments 12 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.05696 2025-10-31 cs.CV 78%

MoralCLIP: Contrastive Alignment of Vision-and-Language Representations with Moral Foundations Theory

Ana Carolina Condez, Diogo Tavares, João Magalhães

机构 * NOVA LINCS, NOVA School of Science and Technology(NOVA LINCS,NOVA科学与技术学院)

专题命中 其他安全 :alignment(title,abstract)

Comments Updated version: corresponds to the ACM MM '25 published paper and includes full appendix material

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.25933 2025-10-31 cs.AI cs.HC cs.LG cs.NE 62%

Humains-Junior: A 3.8B Language Model Achieving GPT-4o-Level Factual Accuracy by Directed Exoskeleton Reasoning

Nissan Yaron, Dan Bystritsky, Ben-Etzion Yaron

机构 * Humains AI Research(Humains人工智能研究) Inpris Ltd(Inpris公司)

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.09391 2025-10-31 cs.CL 57%

Comparing human and LLM politeness strategies in free production

Haoran Zhao, Robert D. Hawkins

机构 * Department of Linguistics University of Washington(语言学系华盛顿大学) Department of Linguistics Stanford University(语言学系斯坦福大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL

Comments 25 pages, 5 figures | EMNLP 2025 camera-ready version

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.26242 2025-10-31 cs.AI 57%

Retrieval Augmented Generation-Enhanced Distributed LLM Agents for Generalizable Traffic Signal Control with Emergency Vehicles

Xinhang Li, Qing Guo, Junyu Chen, Zheng Guo, Shengzhe Xu, Lei Li, Lin Zhang

机构 * School of Artificial Intelligence, Beijing University of Posts(人工智能学院,北京邮电大学) Beijing Big Data Center(北京大数据中心)

专题命中 其他安全 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏