arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-09-17 至 2025-09-17 共收录 34 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 偏好对齐 5 篇

2509.12750 2025-09-17 cs.CV 71%

What Makes a Good Generated Image? Investigating Human and Multimodal LLM Image Preference Alignment

Rishab Parthasarathy, Jasmine Collins, Cory Stephenson

专题命中 偏好对齐 :alignment(title)

Comments 7 pages, 9 figures, 3 tables; appendix 16 pages, 9 figures, 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10105 2025-09-17 cs.CV cs.CL 70%

VARCO-VISION-2.0 Technical Report

Young-rok Cha, Jeongho Ju, SunYoung Park, Jong-Hyeon Lee, Younghyun Yu, Youngjune Kim

机构 * NC AI

专题命中 偏好对齐 :alignment(abstract);safety(abstract);分类 cs.CL

Comments 19 pages, 1 figure, 14 tables. Technical report for VARCO-VISION-2.0, a Korean-English bilingual VLM in 14B and 1.7B variants. Key features: multi-image understanding, OCR with text localization, improved Korean capabilities

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.12521 2025-09-17 cs.LG 57%

Phi: Preference Hijacking in Multi-modal Large Language Models at Inference Time

Yifan Lan, Yuanpu Cao, Weitong Zhang, Lu Lin, Jinghui Chen

机构 * The Pennsylvania State University(宾夕法尼亚州立大学) The University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校)

专题命中 偏好对齐 :safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.12870 2025-09-17 eess.SP 50%

Towards personalized, precise and survey-free environment recognition: AI-enhanced sensor fusion without pre-deployment

Ruichen Wang, Zhikang Ni, Pengzhou Wang, Xiya Cao, Zhi Li, Bao Zhang

专题命中 偏好对齐 :RLHF(abstract)

Comments 5 pages, 7 figures, conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17711 2025-09-17 cs.SI 50%

Enhancing LLM-Based Social Bot via an Adversarial Learning Framework

Fanqi Kong, Xiaoyuan Zhang, Xinyu Chen, Yaodong Yang, Song-Chun Zhu, Xue Feng

专题命中 偏好对齐 :DPO(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 安全训练 3 篇

2509.12060 2025-09-17 cs.AI 83%

When Safe Unimodal Inputs Collide: Optimizing Reasoning Chains for Cross-Modal Safety in Multimodal Large Language Models

Wei Cai, Shujuan Liu, Jian Zhao, Ziyan Shi, Yusheng Zhao, Yuchen Yuan, Tianle Zhang, Chi Zhang, Xuelong Li

专题命中 安全训练 :safety(title,abstract);alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.15805 2025-09-17 cs.CL 70%

Keep Security! Benchmarking Security Policy Preservation in Large Language Model Contexts Against Indirect Attacks in Question Answering

Hwan Chang, Yumin Kim, Yonghyun Jun, Hwanhee Lee

机构 * Chung-Ang University(Chung-Ang 大学)

专题命中 安全训练 :alignment(abstract);safety(abstract);分类 cs.CL

Comments EMNLP 2025 (Main Conference)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.12378 2025-09-17 eess.SY cs.SY 50%

Platoon-Centric Green Light Optimal Speed Advisory Using Safe Reinforcement Learning

Ruining Yang, Jingyuan Zhou, Qiqing Wang, Jinhao Liang, Kaidi Yang

专题命中 安全训练 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 越狱攻击 6 篇

2509.12221 2025-09-17 cs.LG cs.AI cs.CL cs.CR 75%

MEUV: Achieving Fine-Grained Capability Activation in Large Language Models via Mutually Exclusive Unlock Vectors

Xin Tong, Zhi Lin, Jingya Wang, Meng Han, Bo Jin

机构 * Xin Tong School of Information and Cyber Security People’s Public Security University of China(信息与网络安全学院 中国人民公安大学) Zhi Lin School of Safety Science Tsinghua University(安全科学学院 清华大学) Jingya Wang School of Information and Cyber Security People’s Public Security University of China(信息与网络安全学院 中国人民公安大学) Meng Han Intelligent Fusion Research Center Zhejiang University(智能融合研究中心 浙江大学) Bo Jin* The Third Research Institute of the Ministry of Public Security of China(中华人民共和国公安部第三研究所)

专题命中 越狱攻击 :alignment(abstract);safety(abstract);分类 cs.CL、cs.AI、cs.LG

Comments Under Review

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.12937 2025-09-17 cs.CR cs.AI cs.CL 73%

Jailbreaking Large Language Models Through Content Concretization

Johan Wahréus, Ahmed Hussain, Panos Papadimitratos

机构 * Networked Systems Security (NSS) Group(网络系统安全组) KTH Royal Institute of Technology(皇家理工学院)

专题命中 越狱攻击 :safety(abstract);jailbreak(abstract);分类 cs.CL、cs.AI

Comments Accepted for presentation in the Conference on Game Theory and AI for Security (GameSec) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.12724 2025-09-17 cs.CV cs.AI 70%

Defense-to-Attack: Bypassing Weak Defenses Enables Stronger Jailbreaks in Vision-Language Models

Yunhan Zhao, Xiang Zheng, Xingjun Ma

机构 * Fudan University(复旦大学) City University of Hong Kong(香港城市大学)

专题命中 越狱攻击 :safety(abstract);jailbreak(abstract);分类 cs.AI

Comments This work has been submitted to the IEEE for possible publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.13266 2025-09-17 cs.LG cs.AI 62%

JANUS: A Dual-Constraint Generative Framework for Stealthy Node Injection Attacks

Jiahao Zhang, Xiaobing Pei, Zhaokun Zhong, Wenqiang Hao, Zhenghao Tang

专题命中 越狱攻击 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.02844 2025-09-17 cs.CV cs.CL cs.CR 57%

Visual Contextual Attack: Jailbreaking MLLMs with Image-Driven Context Injection

Ziqi Miao, Yi Ding, Lijun Li, Jing Shao

机构 * Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Purdue University(普渡大学)

专题命中 越狱攻击 :jailbreak(abstract);分类 cs.CL

Comments Accepted to EMNLP 2025 (Main). 17 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.15244 2025-09-17 cs.CV cs.AI 57%

Adversarial Prompt Distillation for Vision-Language Models

Lin Luo, Xin Wang, Bojia Zi, Shihao Zhao, Xingjun Ma, Yu-Gang Jiang

机构 * Shanghai Key Lab of Intell. Info. Processing, School of CS, Fudan University(上海智能信息处理实验室,计算机科学学院,复旦大学) The Chinese University of Hong Kong, Shatin, Hong Kong(香港中文大学,沙田,香港) The University of Hong Kong, Pokfulam, Hong Kong(香港大学,薄扶林,香港)

专题命中 越狱攻击 :safety(abstract);分类 cs.AI

Comments This work has been submitted to the IEEE for possible publication

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 提示注入 1 篇

2508.20890 2025-09-17 cs.CR 78%

PromptSleuth: Detecting Prompt Injection via Semantic Intent Invariance

Mengxiao Wang, Yuxuan Zhang, Guofei Gu

专题命中 提示注入 :prompt injection(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 幻觉与事实性 1 篇

2509.12406 2025-09-17 cs.LG stat.ML 57%

Bayesian Parametric Matrix Models: Principled Uncertainty Quantification for Spectral Learning

Mohammad Nooraiepour

专题命中 幻觉与事实性 :safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏

6. 隐私与版权 1 篇

2509.13072 2025-09-17 cs.CR 50%

Digital Sovereignty Control Framework for Military AI-based Cyber Security

Clara Maathuis, Kasper Cools

专题命中 隐私与版权 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

7. 安全评测 11 篇

2509.12936 2025-09-17 cs.LG cs.CL 90%

Rethinking the Evaluation of Alignment Methods: Insights into Diversity, Generalisation, and Safety

Denis Janiak, Julia Moska, Dawid Motyka, Karolina Seweryn, Paweł Walkowiak, Bartosz Żuk, Arkadiusz Janz

机构 * Wroclaw University of Science and Technology (WUST)(沃拉布大学科学与技术学院) National Research Institute (NASK)(国家研究 institute) Institute of Computer Science, Polish Academy of Sciences (IPI PAN)(波兰科学院计算机科学研究所)

专题命中 安全评测 :alignment(title,abstract);safety(title,abstract);DPO(abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.13244 2025-09-17 cs.CL 79%

Evaluating LLM Alignment on Personality Inference from Real-World Interview Data

Jianfeng Zhu, Julina Maharjan, Xinyu Li, Karin G. Coifman, Ruoming Jin

专题命中 安全评测 :alignment(title,abstract);分类 cs.CL

Comments 8 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.12233 2025-09-17 cs.CR cs.AI cs.ET cs.LG cs.NI 76%

Towards Trustworthy Agentic IoEV: AI Agents for Explainable Cyberthreat Mitigation and State Analytics

Meryem Malak Dif, Mouhamed Amine Bouchiha, Abdelaziz Amara Korba, Yacine Ghamri-Doudane

机构 * L3i - La Rochelle University, La Rochelle, France(L3i - 拉罗谢尔大学)

专题命中 安全评测 :trustworthy(title);分类 cs.AI、cs.LG

Comments 10 pages, 7 figures, Accepted at LCN'25

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.12740 2025-09-17 cs.RO cs.AI cs.ET cs.LG cs.SY eess.SY 62%

Deep Generative and Discriminative Digital Twin endowed with Variational Autoencoder for Unsupervised Predictive Thermal Condition Monitoring of Physical Robots in Industry 6.0 and Society 6.0

Eric Guiffo Kaigom

机构 * Department of Computer Science \& Engineering, Frankfurt University of Applied Sciences, Frankfurt a.M., Germany (e-mail: ).

专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.LG

Comments $©$ 2025 the authors. This work has been accepted to the to the 10th IFAC Symposium on Mechatronic Systems & 14th IFAC Symposium on Robotics July 15-18, 2025 || Paris, France for publication under a Creative Commons Licence CC-BY-NC-ND

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.12259 2025-09-17 cs.LG cs.AI quant-ph 62%

Quantum-Inspired Stacked Integrated Concept Graph Model (QISICGM) for Diabetes Risk Prediction

Kenneth G. Young

机构 * II (September 12, 2025)(II)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG

Comments 13 pages, 3 figures, includes performance tables and visualizations. Proposes a Quantum-Inspired Stacked Integrated Concept Graph Model (QISICGM) that integrates phase feature mapping, self-improving concept graphs, and neighborhood sequence modeling within a stacked ensemble. Demonstrates improved F1 and AUC on an augmented PIMA Diabetes dataset with efficient CPU inference

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.12612 2025-09-17 cs.AI 57%

GBV-SQL: Guided Generation and SQL2Text Back-Translation Validation for Multi-Agent Text2SQL

Daojun Chen, Xi Wang, Shenyuan Ren, Qingzhi Ma, Pengpeng Zhao, An Liu

专题命中 安全评测 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.12459 2025-09-17 cs.CL 57%

Does Language Model Understand Language?

Suvojit Acharjee, Utathya Aich, Asfak Ali

机构 * Institute of Engineering and Management(工程与管理研究所) Jadavpur University(贾达沃大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.10483 2025-09-17 cs.SE cs.AI cs.PL 57%

Enhancing Automated Loop Invariant Generation for Complex Programs with Large Language Models

Ruibang Liu, Minyu Chen, Ling-I Wu, Jingyu Ke, Guoqiang Li

机构 * Shanghai Jiao Tong University(上海交通大学)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

Comments 26 pages, 11 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.12912 2025-09-17 cs.RO 50%

Spotting the Unfriendly Robot -- Towards better Metrics for Interactions

Raphael Wenzel, Malte Probst

机构 * Honda Research Institute Europe GmbH(本田欧洲研究院)

专题命中 安全评测 :safety(abstract)

Comments Presented at 2025 IEEE Conference on Robotics and Automation (ICRA) Workshop: Advances in Social Navigation: Planning, HRI and Beyond

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.12492 2025-09-17 cs.CV 50%

Evaluating Robustness of Vision-Language Models Under Noisy Conditions

Purushoth, Alireza

机构 * University of Nevada Reno(内华达大学拉斯维加斯分校)

专题命中 安全评测 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11292 2025-09-17 cs.CV 50%

Leveraging Geometric Priors for Unaligned Scene Change Detection

Ziling Liu, Ziwei Chen, Mingqi Gao, Jinyu Yang, Feng Zheng

机构 * Southern University of Science and Technology(南方科技大学) University of Sheffield(谢菲尔德大学) Spatialtemporal AI(时空AI)

专题命中 安全评测 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

8. 其他安全 6 篇

2509.13282 2025-09-17 cs.CL cs.CV cs.LG 62%

ChartGaze: Enhancing Chart Understanding in LVLMs with Eye-Tracking Guided Attention Refinement

Ali Salamatian, Amirhossein Abaskohi, Wan-Cyuan Fan, Mir Rayat Imtiaz Hossain, Leonid Sigal, Giuseppe Carenini

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.LG

Comments EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.19331 2025-09-17 cs.CV cs.AI cs.CL 62%

Talking to DINO: Bridging Self-Supervised Vision Backbones with Language for Open-Vocabulary Segmentation

Luca Barsellotti, Lorenzo Bianchi, Nicola Messina, Fabio Carrara, Marcella Cornia, Lorenzo Baraldi, Fabrizio Falchi, Rita Cucchiara

机构 * University of Modena and Reggio Emilia(摩德纳和雷吉奥艾米利亚大学) ISTI-CNR(意大利国家研究委员会ISTI) University of Pisa(比萨大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

Comments ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏