arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-10-06 至 2025-10-06 共收录 39 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 偏好对齐 3 篇

2510.02850 2025-10-06 cs.AI 83%

Reward Model Routing in Alignment

Xinle Wu, Yao Lu

机构 * National University of Singapore(新加坡国立大学)

专题命中 偏好对齐 :alignment(title,abstract);RLHF(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03231 2025-10-06 cs.CL cs.AI 62%

Reward Models are Metrics in a Trench Coat

Sebastian Gehrmann

机构 * Bloomberg(高盛)

专题命中 偏好对齐 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02338 2025-10-06 cs.CL cs.AI 62%

Optimizing Long-Form Clinical Text Generation with Claim-Based Rewards

Samyak Jhaveri, Praphul Singh, Jangwon Kim, Tara Taghavi, Krishnaram Kenthapadi

机构 * Oracle Health AI(Oracle健康AI)

专题命中 偏好对齐 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 安全训练 5 篇

2510.02768 2025-10-06 cs.LG cs.CL 81%

A Granular Study of Safety Pretraining under Model Abliteration

Shashank Agnihotri, Jonas Jakubassa, Priyam Dey, Sachin Goyal, Bernt Schiele, Venkatesh Babu Radhakrishnan, Margret Keuper

机构 * Data and Web Science Group, University of Mannheim(曼海姆大学数据与网络科学组) Vision and AI Lab, Indian Institute of Science(印度科学院视觉与人工智能实验室) Carnegie Mellon University(卡内基梅隆大学) Max-Planck-Institute for Informatics, Saarland Informatics Campus(马克斯·普朗克信息研究所,萨尔兰信息校园)

专题命中 安全训练 :safety(title,abstract);分类 cs.CL、cs.LG

Comments Accepted at NeurIPS 2025 bWorkshop Lock-LLM. *Equal Contribution

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02363 2025-10-06 eess.SY cs.SY 78%

Precise HDV Positioning through Safety-Aware Integrated Sensing and Communication in a Value-of-Information-Driven 6G V2X System

Mohammad Reza Abedi, Zahra Rashidi, Nader Mokari, Hamid Saeedi, Nizar Zorba

专题命中 安全训练 :safety(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.17746 2025-10-06 cs.LG cs.AI cs.CL 67%

Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains

Anisha Gunjal, Anthony Wang, Elaine Lau, Vaskar Nath, Yunzhong He, Bing Liu, Sean Hendryx

专题命中 安全训练 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02627 2025-10-06 cs.RO cs.AI 57%

A Trajectory Generator for High-Density Traffic and Diverse Agent-Interaction Scenarios

Ruining Yang, Yi Xu, Yixiao Chen, Yun Fu, Lili Su

机构 * Department of Electrical and Computer Engineering, Northeastern University(电气与计算机工程系,东北大学)

专题命中 安全训练 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02403 2025-10-06 q-bio.QM cs.AI cs.CV 57%

Glaucoma Detection and Structured OCT Report Generation via a Fine-tuned Multimodal Large Language Model

Jalil Jalili, Yashraj Gavhane, Evan Walker, Anna Heinke, Christopher Bowd, Akram Belghith, Massimo A. Fazio, Christopher A. Girkin, C. Gustavo De Moraes, Jeffrey M. Liebmann, Sally L. Baxter, Robert N. Weinreb, Linda M. Zangwill, Mark Christopher

专题命中 安全训练 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 越狱攻击 4 篇

2412.07192 2025-10-06 cs.CR cs.CL cs.LG 79%

PrisonBreak: Jailbreaking Large Language Models with at Most Twenty-Five Targeted Bit-flips

Zachary Coalson, Jeonghyun Woo, Chris S. Lin, Joyce Qu, Yu Sun, Shiyang Chen, Lishan Yang, Gururaj Saileshwar, Prashant Nair, Bo Fang, Sanghyun Hong

机构 * Oregon State University(俄勒冈州立大学) University of British Columbia(不列颠哥伦比亚大学) University of Toronto(多伦多大学) George Mason University(乔治·梅森大学) Rutgers University(罗格斯大学) University of Texas at Arlington(德克萨斯大学阿灵顿分校)

专题命中 越狱攻击 :alignment(abstract);safety(abstract);jailbreak(abstract);分类 cs.CL、cs.LG

Comments Pre-print

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01494 2025-10-06 cs.LG cs.AI 73%

Understanding Adversarial Transfer: Why Representation-Space Attacks Fail Where Data-Space Attacks Succeed

Isha Gupta, Rylan Schaeffer, Joshua Kazdan, Ken Ziyu Liu, Sanmi Koyejo

机构 * ETH Zürich(苏黎世联邦理工学院) Stanford CS(斯坦福大学计算机科学系)

专题命中 越狱攻击 :alignment(abstract);jailbreak(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21634 2025-10-06 cs.CR cs.AI cs.LG cs.NI 73%

MobiLLM: An Agentic AI Framework for Closed-Loop Threat Mitigation in 6G Open RANs

Prakhar Sharma, Haohuang Wen, Vinod Yegneswaran, Ashish Gehani, Phillip Porras, Zhiqiang Lin

机构 * SRI The Ohio State University(俄亥俄州立大学)

专题命中 越狱攻击 :safety(abstract);trustworthy(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03204 2025-10-06 cs.CL 57%

FocusAgent: Simple Yet Effective Ways of Trimming the Large Context of Web Agents

Imene Kerboua, Sahar Omidi Shayegan, Megh Thakkar, Xing Han Lù, Léo Boisvert, Massimo Caccia, Jérémy Espinas, Alexandre Aussem, Véronique Eglin, Alexandre Lacoste

机构 * LIRIS - CNRS, INSA Lyon, Universite Claude Bernard Lyon 1(LIRIS - CNRS,INSA里昂,克劳德·贝尔纳大学里昂) Esker ServiceNow Research(ServiceNow研究) Mila - Quebec AI Institute(魁北克人工智能研究所) McGill University(麦吉尔大学) Polytechnique Montréal(蒙特利尔理工学院)

专题命中 越狱攻击 :prompt injection(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 幻觉与事实性 3 篇

2510.03136 2025-10-06 cs.CL 70%

Beyond the Final Layer: Intermediate Representations for Better Multilingual Calibration in Large Language Models

Ej Zhou, Caiqi Zhang, Tiancheng Hu, Chengzu Li, Nigel Collier, Ivan Vulić, Anna Korhonen

机构 * Language Technology Lab, University of Cambridge(剑桥大学语言技术实验室)

专题命中 幻觉与事实性 :alignment(abstract);trustworthy(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.02580 2025-10-06 cs.AI 70%

V2X-UniPool: Unifying Multimodal Perception and Knowledge Reasoning for Autonomous Driving

Xuewen Luo, Fengze Yang, Fan Ding, Xiangbo Gao, Shuo Xing, Yang Zhou, Zhengzhong Tu, Chenxi Liu

机构 * University of Utah(犹他大学) Monash University(莫纳什大学) Texas A&M University(德克萨斯农工大学)

专题命中 幻觉与事实性 :safety(abstract);trustworthy(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02571 2025-10-06 cs.CV cs.AI cs.CL 62%

How Confident are Video Models? Empowering Video Models to Express their Uncertainty

Zhiting Mei, Ola Shorinwa, Anirudha Majumdar

机构 * Princeton University(普林斯顿大学)

专题命中 幻觉与事实性 :safety(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 隐私与版权 1 篇

2510.02357 2025-10-06 cs.CR cs.AI 57%

Privacy in the Age of AI: A Taxonomy of Data Risks

Grace Billiris, Asif Gill, Madhushi Bandara

专题命中 隐私与版权 :trustworthy(abstract);分类 cs.AI

Comments 12 pages, 2 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏

6. 安全评测 13 篇

2510.02987 2025-10-06 cs.CV 78%

TIT-Score: Evaluating Long-Prompt Based Text-to-Image Alignment via Text-to-Image-to-Text Consistency

Juntong Wang, Huiyu Duan, Jiarui Wang, Ziheng Jia, Guangtao Zhai, Xiongkuo Min

机构 * Institute of Image Communication and Network Engineering(图像通信与网络工程研究所) MoE Key Lab of Artificial Intelligence, AI Institute(人工智能关键实验室) Shanghai Jiao Tong University(上海交通大学)

专题命中 安全评测 :alignment(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02334 2025-10-06 cs.CL cs.AI cs.LG 75%

Where Did It Go Wrong? Attributing Undesirable LLM Behaviors via Representation Gradient Tracing

Zhe Li, Wei Zhao, Yige Li, Jun Sun

机构 * Singapore Management University(新加坡管理大学)

专题命中 安全评测 :safety(abstract);AI safety(abstract);分类 cs.CL、cs.AI、cs.LG

Comments 16 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02677 2025-10-06 cs.AI cs.LG 73%

ARMs: Adaptive Red-Teaming Agent against Multimodal Models with Plug-and-Play Attacks

Zhaorun Chen, Xun Liu, Mintong Kang, Jiawei Zhang, Minzhou Pan, Shuang Yang, Bo Li

机构 * University of Chicago(芝加哥大学) University of Illinois(伊利诺伊大学) Virtue AI Meta

专题命中 安全评测 :alignment(abstract);safety(abstract);分类 cs.AI、cs.LG

Comments 60 pages, 16 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.00038 2025-10-06 cs.LG cs.AI cs.CY 67%

DM-Bench: Benchmarking LLMs for Personalized Decision Making in Diabetes Management

Maria Ana Cardei, Josephine Lamp, Mark Derdzinski, Karan Bhatia

机构 * University of Virginia(弗吉尼亚大学) Dexcom(德科姆公司)

专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.CY、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.00476 2025-10-06 cs.SE cs.AI cs.LG 62%

Analyzing Latent Concepts in Code Language Models

Arushi Sharma, Vedant Pungliya, Christopher J. Quinn, Ali Jannesari

机构 * Iowa State University(爱荷华州立大学)

专题命中 安全评测 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.16076 2025-10-06 cs.CL 57%

The Prompt Makes the Person(a): A Systematic Evaluation of Sociodemographic Persona Prompting for Large Language Models

Marlene Lutz, Indira Sen, Georg Ahnert, Elisa Rogers, Markus Strohmaier

机构 * University of Mannheim(曼海姆大学) GESIS - Leibniz Institute for the Social Sciences(莱布尼茨社会科学研究所) Complexity Science Hub Vienna(维也纳复杂科学中心)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

Comments Accepted to EMNLP Findings 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.02075 2025-10-06 cs.RO cs.LG cs.SY eess.SY 57%

Active Alignments of Lens Systems with Reinforcement Learning

Matthias Burkhardt, Tobias Schmähling, Pascal Stegmann, Michael Layh, Tobias Windisch

机构 * Institute for Machine Vision, University of Applied Sciences Kempten(机器视觉研究所,应用科技大学凯普腾)

专题命中 安全评测 :alignment(abstract);分类 cs.LG

Comments This work has been submitted to the IEEE for possible publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.21185 2025-10-06 cs.LG 57%

Amelia: A Large Dataset and Benchmark for Airport Surface Movement Forecasting

Ingrid Navarro, Pablo Ortega-Kral, Jay Patrikar, Haichuan Wang, Alonso Cano, Zelin Ye, Jong Hoon Park, Sebastian Scherer, Jean Oh

专题命中 安全评测 :safety(abstract);分类 cs.LG

Comments 40 pages, 19 figures, 9 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02994 2025-10-06 cs.CV 50%

Towards Scalable and Consistent 3D Editing

Ruihao Xia, Yang Tang, Pan Zhou

机构 * East China University of Science and Technology(东华大学) Singapore Management University(新加坡管理学院)

专题命中 安全评测 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02909 2025-10-06 cs.CV 50%

Training-Free Out-Of-Distribution Segmentation With Foundation Models

Laith Nayal, Hadi Salloum, Ahmad Taha, Yaroslav Kholodov, Alexander Gasnikov

机构 * Laboratory of Multimodal Research In Industry, AI Institute, Innopolis University(工业多模态研究实验室,人工智能研究所,因诺普利斯大学) Phystech School of Applied Mathematics and Computer Science, Moscow Institute of Physics and Technology(物理与技术莫斯科应用数学与计算机科学学院,莫斯科物理技术学院) Research Center for Artificial Intelligence, Innopolis University(人工智能研究中心,因诺普利斯大学) Q Deep, Innopolis(Q深度,因诺普利斯) Machine Learning and Data Representation Lab, Innopolis University(机器学习与数据表示实验室,因诺普利斯大学)

专题命中 安全评测 :safety(abstract)

Comments 12 pages, 5 figures, 2 tables, ICOMP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02722 2025-10-06 cs.CV 50%

MoGIC: Boosting Motion Generation via Intention Understanding and Visual Context

Junyu Shi, Yong Sun, Zhiyuan Zhang, Lijiang Liu, Zhengjie Zhang, Yuxin He, Qiang Nie

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))

专题命中 安全评测 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20088 2025-10-06 cs.CV cs.MM cs.SD 50%

AudioStory: Generating Long-Form Narrative Audio with Large Language Models

Yuxin Guo, Teng Wang, Yuying Ge, Shijie Ma, Yixiao Ge, Wei Zou, Ying Shan

机构 * School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) ARC Lab, Tencent PCG(腾讯PCG ARC实验室) MAIS, Institute of Automation, CAS, Beijing(自动化研究所北京研究所MAIS)

专题命中 安全评测 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.01534 2025-10-06 cs.CV 50%

Toward a Holistic Evaluation of Robustness in CLIP Models

Weijie Tu, Weijian Deng, Tom Gedeon

专题命中 安全评测 :safety(abstract)

Comments Accepted to IEEE TPAMI, extension of NeurIPS'23 work: A Closer Look at the Robustness of Contrastive Language-Image Pre-Training (CLIP)

详情

展开后加载摘要…

URL PDF HTML 收藏

7. AI治理与伦理 3 篇

2510.03004 2025-10-06 cs.LG cs.AI 62%

BrainIB++: Leveraging Graph Neural Networks and Information Bottleneck for Functional Brain Biomarkers in Schizophrenia

Tianzheng Hu, Qiang Li, Shu Liu, Vince D. Calhoun, Guido van Wingen, Shujian Yu

机构 * Vrije University Amsterdam(荷兰阿姆斯特丹自由大学) Tri-institutional Center for Translational Research in Neuroimaging(转化神经影像研究联合中心) Emory University(埃默里大学) Key Laboratory of Genetic Evolution and Animal Models(遗传进化与动物模型重点实验室) Kunming Institute of Zoology(昆明动物研究所) Chinese Academy of Sciences Kunming(中国科学院昆明分院) Department of Psychiatry, Amsterdam UMC, University of Amsterdam(阿姆斯特丹大学精神病科) Department of Physics and Technology, UiT The Arctic University of Norway(北极大学挪威理工学院物理与技术系)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.LG

Comments This manuscript has been accepted by Biomedical Signal Processing and Control and the code is available at https://github.com/TianzhengHU/BrainIB_coding/tree/main/BrainIB_GIB

详情

展开后加载摘要…

URL PDF HTML 收藏