arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-07-25 至 2025-07-25 共收录 18 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 偏好对齐 2 篇

2502.03699 2025-07-25 cs.CL cs.AI cs.IR 84%

LLM Alignment as Retriever Optimization: An Information Retrieval Perspective

Bowen Jin, Jinsung Yoon, Zhen Qin, Ziqi Wang, Wei Xiong, Yu Meng, Jiawei Han, Sercan O. Arik

机构 * University of Illinois at Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Google Cloud AI Research(谷歌云人工智能研究) Google DeepMind(谷歌DeepMind) University of Virginia(弗吉尼亚大学)

专题命中 偏好对齐 :alignment(title,abstract);trustworthy(abstract);分类 cs.CL、cs.AI

Comments 26 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.18417 2025-07-25 cs.CL cs.LG q-fin.ST q-fin.TR 73%

FinDPO: Financial Sentiment Analysis for Algorithmic Trading through Preference Optimization of LLMs

Giorgos Iacovides, Wuyang Zhou, Danilo Mandic

机构 * Imperial College London(伦敦帝国理工学院)

专题命中 偏好对齐 :alignment(abstract);DPO(abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 安全训练 2 篇

2506.23276 2025-07-25 cs.AI cs.CL 62%

Corrupted by Reasoning: Reasoning Language Models Become Free-Riders in Public Goods Games

David Guzman Piedrahita, Yongjin Yang, Mrinmaya Sachan, Giorgia Ramponi, Bernhard Schölkopf, Zhijing Jin

专题命中 安全训练 :alignment(abstract);分类 cs.CL、cs.AI

Comments Published at COLM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.17868 2025-07-25 eess.SY cs.SY 50%

Safe Reinforcement Learning-based Automatic Generation Control

Amr S. Mohamed, Emily Nguyen, Deepa Kundur

专题命中 安全训练 :safety(abstract)

Comments 5 pages, conference: IEEE Power and Energy Systems General Meeting 2025

Journal ref Mohamed, Amr, Emily Nguyen, and Deepa Kundur. "Safe Reinforcement Learning-based Automatic Generation Control." 2025 IEEE Power & Energy Society General Meeting (PESGM). IEEE, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 越狱攻击 3 篇

2505.02581 2025-07-25 cs.AI 79%

Neurodivergent Influenceability as a Contingent Solution to the AI Alignment Problem

Alberto Hernández-Espinosa, Felipe S. Abrahão, Olaf Witkowski, Hector Zenil

机构 * Oxford Immune Algorithmics(牛津免疫算法公司) Oxford University Innovation(牛津大学创新) London Institute for Healthcare Engineering(伦敦医疗工程研究所) University of Tokyo(东京大学) The Arrival Institute(抵达研究所) Cancer Research Group(癌症研究组) The Francis Crick Institute(弗朗西斯·克里克研究所) Cross Labs(交叉实验室) King’s Institute for Artificial Intelligence(国王人工智能研究所) King’s College London(伦敦国王学院) The Alan Turing Institute(艾伦·图灵研究所)

专题命中 越狱攻击 :alignment(title,abstract);分类 cs.AI

Comments 44 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.17256 2025-07-25 cs.CL 70%

Weak-to-Strong Jailbreaking on Large Language Models

Xuandong Zhao, Xianjun Yang, Tianyu Pang, Chao Du, Lei Li, Yu-Xiang Wang, William Yang Wang

机构 * Sea AI Lab, Singapore(新加坡海智实验室) Carnegie Mellon University(卡内基梅隆大学)

专题命中 越狱攻击 :safety(abstract);jailbreak(abstract);分类 cs.CL

Comments ICML 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.17922 2025-07-25 cs.LG cs.AI 62%

From Seed to Harvest: Augmenting Human Creativity with AI for Red-teaming Text-to-Image Models

Jessica Quaye, Charvi Rastogi, Alicia Parrish, Oana Inel, Minsuk Kahng, Lora Aroyo, Vijay Janapa Reddi

专题命中 越狱攻击 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 幻觉与事实性 1 篇

2507.14660 2025-07-25 cs.AI cs.CL 73%

When Autonomy Goes Rogue: Preparing for Risks of Multi-Agent Collusion in Social Systems

Qibing Ren, Sitao Xie, Longxuan Wei, Zhenfei Yin, Junchi Yan, Lizhuang Ma, Jing Shao

专题命中 幻觉与事实性 :safety(abstract);AI safety(abstract);分类 cs.CL、cs.AI

Comments Code is available at https://github.com/renqibing/MultiAgent4Collusion

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 安全评测 3 篇

2502.04757 2025-07-25 cs.CV cs.CL 83%

ELITE: Enhanced Language-Image Toxicity Evaluation for Safety

Wonjun Lee, Doehyeon Lee, Eugene Choi, Sangyoon Yu, Ashkan Yousefpour, Haon Park, Bumsub Ham, Suhyun Kim

机构 * Yonsei University(延世大学) Kyung Hee University(庆熙大学) Korea Institute of Science(韩国科学研究院) Seoul National University(首尔国立大学) Sookmyung Women's University(_sookmyung女子大学)

专题命中 安全评测 :safety(title,abstract);alignment(abstract);分类 cs.CL

Comments ICML 2025. Project page at https://velpegor.github.io/ELITE/

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.18447 2025-07-25 cs.CV 50%

PDB-Eval: An Evaluation of Large Multimodal Models for Description and Explanation of Personalized Driving Behavior

Junda Wu, Jessica Echterhoff, Kyungtae Han, Amr Abdelraouf, Rohit Gupta, Julian McAuley

机构 * Computer Science and Engineering University of California San Diego(计算机科学与工程大学加州大学圣地亚哥分校) InfoTech Labs Toyota Motor North America(信息科技实验室丰田北美)

专题命中 安全评测 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.18105 2025-07-25 cs.SE cs.CR 50%

Understanding the Supply Chain and Risks of Large Language Model Applications

Yujie Ma, Lili Quan, Xiaofei Xie, Qiang Hu, Jiongchi Yu, Yao Zhang, Sen Chen

专题命中 安全评测 :trustworthy(abstract)

Comments 26 pages

详情

展开后加载摘要…

URL PDF HTML 收藏

6. AI治理与伦理 4 篇

2507.00566 2025-07-25 cs.CV 78%

Zero-Shot Skeleton-Based Action Recognition With Prototype-Guided Feature Alignment

Kai Zhou, Shuhai Zhang, Zeng You, Jinwu Hu, Mingkui Tan, Fei Liu

机构 * School of Software Engineering, South China University of Technology(南方科技大学软件工程学院) South China University of Technology(南方科技大学) Pazhou Lab(琶洲实验室) School of Future Technology, South China University of Technology(未来技术学院) Peng Cheng Laboratory(鹏城实验室) Key Laboratory of Big Data and Intelligent Robot (South China University of Technology), Ministry of Education(大数据与智能机器人重点实验室)

专题命中 AI治理与伦理 :alignment(title,abstract)

Comments This paper is accepted by IEEE TIP 2025 (The journal version is available at https://doi.org/10.1109/TIP.2025.3586487). Code is publicly available at https://github.com/kaai520/PGFA

Journal ref IEEE Transactions on Image Processing 34 (2025) 4602-4617

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.17788 2025-07-25 cs.LG cs.AI 62%

Adaptive Repetition for Mitigating Position Bias in LLM-Based Ranking

Ali Vardasbi, Gustavo Penha, Claudia Hauff, Hugues Bouchard

机构 * Spotify

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.17787 2025-07-25 cs.LG cs.AI 62%

Hyperbolic Deep Learning for Foundation Models: A Survey

Neil He, Hiren Madhu, Ngoc Bui, Menglin Yang, Rex Ying

机构 * Yale University(耶鲁大学) Hong Kong University of Science(香港科学大学)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.LG

Comments 11 Pages, SIGKDD 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.18328 2025-07-25 cs.NI 50%

Enhanced Velocity-Adaptive Scheme: Joint Fair Access and Age of Information Optimization in Vehicular Networks

Xiao Xu, Qiong Wu, Pingyi Fan, Kezhi Wang, Nan Cheng, Wen Chen, Khaled B. Letaief

专题命中 AI治理与伦理 :safety(abstract)

Comments This paper has been submitted to IEEE TMC

详情

展开后加载摘要…

URL PDF HTML 收藏

7. 其他安全 3 篇

2507.16680 2025-07-25 cs.LG cs.IT cs.NI math.IT 74%

Latent Space Alignment for AI-Native MIMO Semantic Communications

Mario Edoardo Pandolfo, Simone Fiorellino, Emilio Calvanese Strinati, Paolo Di Lorenzo

机构 * DIAG Department, Sapienza University of Rome(罗马萨皮恩扎大学DIAG系) Consorzio Nazionale Interuniversitario per le Telecomunicazioni (CNIT)(国家跨大学电信合作组织(CNIT)) CEA Leti, University Grenoble Alpes(CEA Leti,格勒诺布尔阿尔卑斯大学) DIET Department, Sapienza University of Rome(罗马萨皮恩扎大学DIET系)

专题命中 其他安全 :alignment(title);分类 cs.LG

Comments Proc. of IEEE IJCNN 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.17874 2025-07-25 cs.AI 57%

I2I-STRADA -- Information to Insights via Structured Reasoning Agent for Data Analysis

SaiBarath Sundar, Pranav Satheesan, Udayaadithya Avadhanam

机构 * Mphasis Limited(默比斯有限公司)

专题命中 其他安全 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.23825 2025-07-25 cs.CV 50%

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams

Haoji Zhang, Yiqin Wang, Yansong Tang, Yong Liu, Jiashi Feng, Xiaojie Jin

机构 * Tsinghua University(清华大学) Tsinghua Shenzhen International Graduate School(清华大学深圳国际研究生院) Beijing Jiaotong University(北京交通大学) ByteDance Inc(字节跳动公司)

专题命中 其他安全 :alignment(abstract)

Comments Accepted by ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏