arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-11-07 至 2025-11-07 共收录 41 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 偏好对齐 6 篇

2511.03939 2025-11-07 cs.LG cs.AI cs.CL 90%

RLHF: A comprehensive Survey for Cultural, Multimodal and Low Latency Alignment Methods

Raghav Sharma, Manan Mehta, Sai Tiger Raina

机构 * Northeastern University(东北大学) University of Southern California(南加州大学)

专题命中 偏好对齐 :alignment(title,abstract);RLHF(title,abstract);DPO(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.03997 2025-11-07 cs.CV 82%

PhysCorr: Dual-Reward DPO for Physics-Constrained Text-to-Video Generation with Automated Preference Selection

Peiyao Wang, Weining Wang, Qi Li

专题命中 偏好对齐 :DPO(title);alignment(abstract);trustworthy(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.13518 2025-11-07 cs.CL cs.AI cs.LG 75%

Selective Preference Optimization via Token-Level Reward Function Estimation

Kailai Yang, Zhiwei Liu, Qianqian Xie, Jimin Huang, Erxue Min, Sophia Ananiadou

机构 * National Centre for Text Mining, The University of Manchester(曼彻斯特大学文本挖掘中心) School of Artificial Intelligence, Wuhan University(武汉大学人工智能学院) Center for Language and Information Research, Wuhan University(武汉大学语言与信息研究中心) The Fin AI(Fin AI公司) Baidu Inc.(百度公司) Archimedes Research(阿基米德研究)

专题命中 偏好对齐 :alignment(abstract);DPO(abstract);分类 cs.CL、cs.AI、cs.LG

Comments Accepted by the EMNLP 2025 main conference

Journal ref https://aclanthology.org/2025.emnlp-main.359/

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.03823 2025-11-07 cs.CL cs.AI 73%

PLLuM: A Family of Polish Large Language Models

Jan Kocoń, Maciej Piasecki, Arkadiusz Janz, Teddy Ferdinan, Łukasz Radliński, Bartłomiej Koptyra, Marcin Oleksy, Stanisław Woźniak, Paweł Walkowiak, Konrad Wojtasik, Julia Moska, Tomasz Naskręt, Bartosz Walkowiak, Mateusz Gniewkowski, Kamil Szyc, Dawid Motyka, Dawid Banach, Jonatan Dalasiński, Ewa Rudnicka, Bartłomiej Alberski, Tomasz Walkowiak, Aleksander Szczęsny, Maciej Markiewicz, Tomasz Bernaś, Hubert Mazur, Kamil Żyta, Mateusz Tykierko, Grzegorz Chodak, Tomasz Kajdanowicz, Przemysław Kazienko, Agnieszka Karlińska, Karolina Seweryn, Anna Kołos, Maciej Chrabąszcz, Katarzyna Lorenc, Aleksandra Krasnodębska, Artur Wilczek, Katarzyna Dziewulska, Paula Betscher, Zofia Cieślińska, Katarzyna Kowol, Daria Mikoś, Maciej Trzciński, Dawid Krutul, Marek Kozłowski, Sławomir Dadas, Rafał Poświata, Michał Perełkiewicz, Małgorzata Grębowiec, Maciej Kazuła, Marcin Białas, Roman Roszko, Danuta Roszko, Jurgita Vaičenonienė, Andrius Utka, Paweł Levchuk, Paweł Kowalski, Irena Prawdzic-Jankowska, Maciej Ogrodniczuk, Monika Borys, Anna Bulińska, Wiktoria Gumienna, Witold Kieraś, Dorota Komosińska, Katarzyna Krasnowska-Kieraś, Łukasz Kobyliński, Martyna Lewandowska, Marek Łaziński, Mikołaj Łątkowski, Dawid Mastalerz, Beata Milewicz, Agnieszka Anna Mykowiecka, Angelika Peljak-Łapińska, Sandra Penno, Zuzanna Przybysz, Michał Rudolf, Piotr Rybak, Karolina Saputa, Aleksandra Tomaszewska, Aleksander Wawer, Marcin Woliński, Joanna Wołoszyn, Alina Wróblewska, Bartosz Żuk, Filip Żarnecki, Konrad Kaczyński, Anna Cichosz, Zuzanna Deckert, Monika Garnys, Izabela Grabarczyk, Wojciech Janowski, Sylwia Karasińska, Aleksandra Kujawiak, Piotr Misztela, Maria Szymańska, Karolina Walkusz, Igor Siek, Jakub Kwiatkowski, Piotr Pęzik

机构 * Department of Artificial Intelligence, Wrocław University of Science and Technology(沃斯克拉大学人工智能系) NASK National Research Institute(国家研究 institute) National Information Processing Institute(国家信息处理研究所) Institute of Slavic Studies, Polish Academy of Sciences(波兰科学院斯拉夫研究学院) Institute of Computer Science, Polish Academy of Sciences(波兰科学院计算机科学研究所) University of Łódź(卢布林大学)

专题命中 偏好对齐 :alignment(abstract);safety(abstract);分类 cs.CL、cs.AI

Comments 83 pages, 19 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.13444 2025-11-07 econ.GN q-fin.EC 71%

Balancing Engagement and Polarization: Multi-Objective Alignment of News Content Using LLMs

Mengjie Cheng, Elie Ofek, Hema Yoganarasimhan

专题命中 偏好对齐 :alignment(title)

Comments 73 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.04286 2025-11-07 cs.LG cs.AI 62%

Efficient Reinforcement Learning from Human Feedback via Bayesian Preference Inference

Matteo Cercola, Valeria Capretti, Simone Formentin

专题命中 偏好对齐 :RLHF(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 安全训练 2 篇

2511.04215 2025-11-07 cs.CR cs.CL 57%

Black-Box Guardrail Reverse-engineering Attack

Hongwei Yao, Yun Xia, Shuo Shao, Haoran Shi, Tong Qiao, Cong Wang

机构 * City University of Hong Kong(香港城市大学) Zhejiang University(浙江大学) Hangzhou Dianzi University(杭州电子科技大学)

专题命中 安全训练 :safety(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.04147 2025-11-07 cs.LG 57%

Exchange Policy Optimization Algorithm for Semi-Infinite Safe Reinforcement Learning

Jiaming Zhang, Yujie Yang, Haoning Wang, Liping Zhang, Shengbo Eben Li

机构 * Department of Mathematical Sciences Tsinghua University(清华大学数学科学系) School of Vehicle and Mobility Tsinghua University(清华大学车辆与移动系统学院) Department of Mathematical Sciences, Tsinghua University(清华大学数学科学系) School of Vehicle and Mobility & College of AI Tsinghua University(清华大学车辆与移动系统学院与人工智能学院)

专题命中 安全训练 :safety(abstract);分类 cs.LG

Comments Submitted to the Journal of Machine Learning Research (JMLR), under review

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 越狱攻击 1 篇

2410.23558 2025-11-07 cs.CR cs.AI 57%

Transferable & Stealthy Ensemble Attacks: A Black-Box Jailbreaking Framework for Large Language Models

Yiqi Yang, Hongye Fu

专题命中 越狱攻击 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 提示注入 1 篇

2511.04508 2025-11-07 cs.CR 50%

Large Language Models for Cyber Security

Raunak Somani, Aswani Kumar Cherukuri

专题命中 提示注入 :prompt injection(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 幻觉与事实性 1 篇

2505.04847 2025-11-07 cs.CL cs.AI 62%

Benchmarking LLM Faithfulness in RAG with Evolving Leaderboards

Manveer Singh Tamber, Forrest Sheng Bao, Chenyu Xu, Ge Luo, Suleman Kazi, Minseok Bae, Miaoran Li, Ofer Mendelevitch, Renyi Qu, Jimmy Lin

机构 * University of Waterloo(滑铁卢大学) Vectara(Vectara公司) Iowa State University(爱荷华州立大学) Stanford University(斯坦福大学)

专题命中 幻觉与事实性 :trustworthy(abstract);分类 cs.CL、cs.AI

Comments EMNLP Industry Track 2025

详情

展开后加载摘要…

URL PDF HTML 收藏

6. 隐私与版权 1 篇

2511.04228 2025-11-07 cs.CL cs.LG 62%

REMIND: Input Loss Landscapes Reveal Residual Memorization in Post-Unlearning LLMs

Liran Cohen, Yaniv Nemcovesky, Avi Mendelson

机构 * Technion - Israel Institute of Technology(技术离子-以色列理工学院)

专题命中 隐私与版权 :safety(abstract);分类 cs.CL、cs.LG

Comments Pre-print version under review

详情

展开后加载摘要…

URL PDF HTML 收藏

7. 安全评测 14 篇

2511.04328 2025-11-07 cs.AI 83%

RxSafeBench: Identifying Medication Safety Issues of Large Language Models in Simulated Consultation

Jiahao Zhao, Luxin Xu, Minghuan Tan, Lichao Zhang, Ahmadreza Argha, Hamid Alinejad-Rokny, Min Yang

机构 * Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences(深圳先进技术研究院,中国科学院) University of Electronic Science and Technology of China(电子科技大学) Shenzhen University of Advanced Technology(深圳大学) School of Biomedical Engineering, UNSW Sydney(生物医学工程学院,UNSW悉尼)

专题命中 安全评测 :safety(title,abstract);trustworthy(abstract);分类 cs.AI

Comments To appear in BIBM2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.07416 2025-11-07 cs.CV cs.CL cs.LG 81%

RadZero: Similarity-Based Cross-Attention for Explainable Vision-Language Alignment in Chest X-ray with Zero-Shot Multi-Task Capability

Jonggwon Park, Byungmu Yoon, Soobum Kim, Kyoyun Choi

专题命中 安全评测 :alignment(title,abstract);分类 cs.CL、cs.LG

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.04312 2025-11-07 cs.AI 79%

Probing the Probes: Methods and Metrics for Concept Alignment

Jacob Lysnæs-Larsen, Marte Eggen, Inga Strümke

机构 * Department of Computer Science NTNU - Norwegian University of Science and Technology(计算机科学系挪威科学技术大学)

专题命中 安全评测 :alignment(title,abstract);分类 cs.AI

Comments 29 pages, 17 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.04184 2025-11-07 cs.CL cs.AI 76%

Trustworthy LLM-Mediated Communication: Evaluating Information Fidelity in LLM as a Communicator (LAAC) Framework in Multiple Application Domains

Mohammed Musthafa Rafi, Adarsh Krishnamurthy, Aditya Balu

机构 * Iowa State University(爱荷华州立大学)

专题命中 安全评测 :trustworthy(title);分类 cs.CL、cs.AI

Comments 10 pages, 4 figures. Submitted to IEEE DISTILL 2025 (co-located with IEEE TPS 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.04502 2025-11-07 cs.CL cs.AI 73%

RAGalyst: Automated Human-Aligned Agentic Evaluation for Domain-Specific RAG

Joshua Gao, Quoc Huy Pham, Subin Varghese, Silwal Saurav, Vedhus Hoskere

机构 * University of Houston(德克萨斯大学休斯敦分校)

专题命中 安全评测 :alignment(abstract);safety(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.04316 2025-11-07 cs.AI cs.SE 70%

AdversariaLLM: A Unified and Modular Toolbox for LLM Robustness Research

Tim Beyer, Jonas Dornbusch, Jakob Steimle, Moritz Ladenburger, Leo Schwinn, Stephan Günnemann

机构 * Department of Computer Science, Technical University of Munich, Germany(慕尼黑技术大学计算机科学系) Munich Data Science Institute, Germany(慕尼黑数据科学研究所)

专题命中 安全评测 :safety(abstract);jailbreak(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.03945 2025-11-07 cs.CL cs.AI 62%

Direct Semantic Communication Between Large Language Models via Vector Translation

Fu-Chun Yang, Jason Eshraghian

机构 * University of California, Santa Cruz(加州大学圣克ruz分校)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

Comments 9 pages, 1 figure, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.04538 2025-11-07 cs.CL 57%

From Model to Breach: Towards Actionable LLM-Generated Vulnerabilities Reporting

Cyril Vallez, Alexander Sternfeld, Andrei Kucharavy, Ljiljana Dolamic

机构 * IEM, HES-SO Valais-Wallis(IEM,HES-SO瓦莱-达沃斯) II, HES-SO Valais-Wallis(II,HES-SO瓦莱-达沃斯) Cyber-Defence Campus, armasuisse(网络安全防御校区,armasuisse)

专题命中 安全评测 :safety(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.06991 2025-11-07 cs.AI cs.GT cs.HC 57%

Evaluating LLM-Contaminated Crowdsourcing Data Without Ground Truth

Yichi Zhang, Jinlong Pang, Zhaowei Zhu, Yang Liu

机构 * DIMACS, Rutgers University(Rutgers大学DIMACS研究中心)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

Comments 32 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.17793 2025-11-07 cs.CL 57%

Compression Hacking: A Supplementary Perspective on Informatics Properties of Language Models from Geometric Distortion

Jianxiang Zang, Meiling Ning, Yongda Wei, Shihan Dou, Jiazheng Zhang, Nijia Mo, Binhong Li, Tao Gui, Qi Zhang, Xuanjing Huang

机构 * Computation and Artificial Intelligence Innovative College, Fudan University(复旦大学计算与人工智能创新学院) Beijing University of Posts and Telecommunications(北京邮电大学) Shanghai University of International Business and Economics(上海国际商务经济学院) Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))

专题命中 安全评测 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.05712 2025-11-07 cs.LG cs.CV q-bio.NC 57%

Scaling Laws for Task-Optimized Models of the Primate Visual Ventral Stream

Abdulkadir Gokce, Martin Schrimpf

机构 * EPFL(苏黎世联邦理工学院)

专题命中 安全评测 :alignment(abstract);分类 cs.LG

Comments Published at ICML25 as a spotlight paper - 9 pages for the main paper, 22 pages in total. 7 main figures and 7 supplementary figures. Code, model weights, and benchmark results can be accessed at https://github.com/epflneuroailab/scaling-primate-vvs

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.13406 2025-11-07 cs.AI cs.CE cs.MA 57%

Collaboration Dynamics and Reliability Challenges of Multi-Agent LLM Systems in Finite Element Analysis

Chuan Tian, Yilei Zhang

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.04012 2025-11-07 cs.SE 50%

PSD2Code: Automated Front-End Code Generation from Design Files via Multimodal Large Language Models

Yongxi Chen, Lei Chen

专题命中 安全评测 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.04267 2025-11-07 cs.SE 50%

A Tool for Benchmarking Large Language Models' Robustness in Assessing the Realism of Driving Scenarios

Jiahui Wu, Chengjie Lu, Aitor Arrieta, Shaukat Ali

专题命中 安全评测 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

8. AI治理与伦理 3 篇

2511.04157 2025-11-07 cs.SE cs.AI 83%

Are We Aligned? A Preliminary Investigation of the Alignment of Responsible AI Values between LLMs and Human Judgment

Asma Yamani, Malak Baslyman, Moataz Ahmed

机构 * Information and Computer Science Department, KFUPM(信息与计算机科学系,KFUPM) IRC for finance and digital economy, KFUPM(金融与数字经济研究中心,KFUPM) SDAIA-KFUPM Joint Research Center for Artificial Intelligence, KFUPM(人工智能联合研究中心,KFUPM)

专题命中 AI治理与伦理 :alignment(title,abstract);safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.02895 2025-11-07 cs.CY cs.AI cs.HC physics.soc-ph 73%

A Criminology of Machines

Gian Maria Campedelli

机构 * Fondazione Bruno Kessler(布鲁诺·科斯勒基金会)

专题命中 AI治理与伦理 :safety(abstract);AI safety(abstract);分类 cs.AI、cs.CY

Comments This pre-print is also available at CrimRxiv with DOI: https://doi.org/10.21428/cb6ab371.e3354ce1

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.03980 2025-11-07 cs.AI cs.CL 62%

LLMs and Cultural Values: the Impact of Prompt Language and Explicit Cultural Framing

Bram Bulté, Ayla Rigouts Terryn

机构 * Brussels Centre for Language Studies, Vrije Universiteit Brussel(布鲁塞尔语言研究中心,布鲁塞尔自由大学) Université de Montréal & Mila - Quebec AI Institute(蒙特利尔大学及魁北克人工智能研究所)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI

Comments Preprint under review at Computational Linguistics. Accepted with minor revisions (10/10/2025); second round

详情

展开后加载摘要…

URL PDF HTML 收藏

9. 其他安全 12 篇

2511.04601 2025-11-07 cs.CV cs.MM 78%

PixCLIP: Achieving Fine-grained Visual Language Understanding via Any-granularity Pixel-Text Alignment Learning

Yicheng Xiao, Yu Chen, Haoxuan Ma, Jiale Hong, Caorui Li, Lingxiang Wu, Haiyun Guo, Jinqiao Wang

专题命中 其他安全 :alignment(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏