arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-08-15 至 2025-08-15 共收录 40 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 偏好对齐 3 篇

2508.10116 2025-08-15 cs.IR 78%

Bridging Modality Gaps in e-Commerce Products via Vision-Language Alignment

Yipeng Zhang, Hongju Yu, Aritra Mandal, Canran Xu, Qunzhi Zhou, Zhe Wu

专题命中 偏好对齐 :alignment(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10701 2025-08-15 cs.LG cs.AI 62%

REFN: A Reinforcement-Learning-From-Network Framework against 1-day/n-day Exploitations

Tianlong Yu, Lihong Liu, Ziyi Zhou, Fudu Xing, Kailong Wang, Yang Yang

机构 * School of Artificial Intelligence, Hubei University(湖北大学人工智能学院) Huazhong University of Science and Technology(华中科技大学) University of Southern California(美国南加州大学)

专题命中 偏好对齐 :RLHF(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.14493 2025-08-15 cs.IR cs.AI cs.LG 62%

FinSage: A Multi-aspect RAG System for Financial Filings Question Answering

Xinyu Wang, Jijun Chi, Zhenghan Tai, Tung Sum Thomas Kwok, Muzhi Li, Zhuhong Li, Hailin He, Yuchen Hua, Peng Lu, Suyuchen Wang, Yihong Wu, Jerry Huang, Jingrui Tian, Fengran Mo, Yufei Cui, Ling Zhou

机构 * 1SimpleWay.AI 2McGill University 3University of Toronto 4University of California, Los Angeles 5The Chinese University of Hong Kong 6Duke University 7Universit\'e de Montr\'eal 8Mila - Quebec AI Institute 9Noah's Ark Lab 10CG Matrix Technology Limited 1SimpleWay.AI 2McGill University 3University of Toronto 4University of California, Los Angeles 5The Chinese University of Hong Kong 6Duke University 7Universit\'e de Montr\'eal 8Mila - Quebec AI Institute 9Noah's Ark Lab 10CG Matrix Technology Limited

专题命中 偏好对齐 :DPO(abstract);分类 cs.AI、cs.LG

Comments Accepted at the 34th ACM International Conference on Information and Knowledge Management (CIKM2025)

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 安全训练 6 篇

2508.10033 2025-08-15 cs.CR cs.AI 70%

Cognitive Cybersecurity for Artificial Intelligence: Guardrail Engineering with CCS-7

Yuksel Aydin

机构 * Independent Researcher(独立研究者)

专题命中 安全训练 :alignment(abstract);safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10355 2025-08-15 cs.CL 57%

Making Qwen3 Think in Korean with Reinforcement Learning

Jungyup Lee, Jemin Kim, Sang Park, SeungJae Lee

机构 * Jungyup Lee, Jemin Kim, Sang Park, SeungJae Lee(作者)

专题命中 安全训练 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10233 2025-08-15 cs.LG 57%

Interpretable Machine Learning Model for Early Prediction of Acute Kidney Injury in Critically Ill Patients with Cirrhosis: A Retrospective Study

Li Sun, Shuheng Chen, Junyi Fan, Yong Si, Minoo Ahmadi, Elham Pishgar, Kamiar Alaei, Maryam Pishgar

专题命中 安全训练 :alignment(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10378 2025-08-15 cs.RO 50%

A Semantic-Aware Framework for Safe and Intent-Integrative Assistance in Upper-Limb Exoskeletons

Yu Chen, Shu Miao, Chunyu Wu, Jingsong Mu, Bo OuYang, Xiang Li

机构 * Department of Automation, Tsinghua University(自动化系,清华大学)

专题命中 安全训练 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01272 2025-08-15 cs.CV 50%

PromptSafe: Gated Prompt Tuning for Safe Text-to-Image Generation

Zonglei Jing, Xiao Yang, Xiaoqian Li, Siyuan Liang, Aishan Liu, Mingchuan Zhang, Xianglong Liu

机构 * Beihang University(北航) Beijing University of Posts and Telecommunications(北京邮电大学) Taishan University(泰山大学) Nanyang Technological University(南洋理工大学) Henan University of Science and Technology(河南科技大学)

专题命中 安全训练 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.00334 2025-08-15 cs.RO 50%

Traversability analysis with vision and terrain probing for safe legged robot navigation

Garen Haddeler, Meng Yee Michael Chuah, Yangwei You, Jianle Chan, Albertus H. Adiwahono, Wei Yun Yau, Chee-Meng Chew

专题命中 安全训练 :safety(abstract)

Journal ref Frontiers in Robotics and AI, Volume 9 - 2022

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 越狱攻击 2 篇

2508.10032 2025-08-15 cs.CL cs.AI 81%

The Cost of Thinking: Increased Jailbreak Risk in Large Language Models

Fan Yang

机构 * Fan Yang

专题命中 越狱攻击 :jailbreak(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10404 2025-08-15 cs.CL cs.AI 79%

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation

Huizhen Shu, Xuying Li, Qirui Wang, Yuji Kosuga, Mengqiu Tian, Zhuo Li

专题命中 越狱攻击 :alignment(abstract);safety(abstract);jailbreak(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 红队测试 1 篇

2508.10108 2025-08-15 cs.AI cs.CL 82%

Amazon Nova AI Challenge -- Trusted AI: Advancing secure, AI-assisted software development

Sattvik Sahai, Prasoon Goyal, Michael Johnston, Anna Gottardi, Yao Lu, Lucy Hu, Luke Dai, Shaohua Liu, Samyuth Sagi, Hangjie Shi, Desheng Zhang, Lavina Vaz, Leslie Ball, Maureen Murray, Rahul Gupta, Shankar Ananthakrishna

机构 * Amazon(亚马逊)

专题命中 红队测试 :alignment(abstract);safety(abstract);red teaming(abstract);AI safety(abstract)

Comments 18 pages, 1st Proceedings of Amazon Nova AI Challenge (Trusted AI 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 幻觉与事实性 3 篇

2508.10010 2025-08-15 cs.CL 57%

An Audit and Analysis of LLM-Assisted Health Misinformation Jailbreaks Against LLMs

Ayana Hussain, Patrick Zhao, Nicholas Vincent

专题命中 幻觉与事实性 :jailbreak(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.09458 2025-08-15 cs.HC cs.AI cs.ET 57%

Hallucination vs interpretation: rethinking accuracy and precision in AI-assisted data extraction for knowledge synthesis

Xi Long, Christy Boscardin, Lauren A. Maggio, Joseph A. Costello, Ralph Gonzales, Rasmyah Hammoudeh, Ki Lai, Yoon Soo Park, Brian C. Gin

专题命中 幻觉与事实性 :trustworthy(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.06776 2025-08-15 cs.RO cs.SY eess.SY 50%

Chance-constrained Linear Quadratic Gaussian Games for Multi-robot Interaction under Uncertainty

Kai Ren, Giulio Salizzoni, Mustafa Emre Gürsoy, Maryam Kamgarpour

机构 * SYCAMORE Lab, École Polytechnique Fédérale de Lausanne (EPFL)(SYCAMORE实验室,瑞士联邦理工学院(EPFL))

专题命中 幻觉与事实性 :safety(abstract)

Comments Published in IEEE Control Systems Letters

Journal ref IEEE Control Systems Letters, vol. 9, pp. 2061-2066, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏

6. 隐私与版权 1 篇

2508.10020 2025-08-15 cs.CL cs.AI 62%

FedCoT: Communication-Efficient Federated Reasoning Enhancement for Large Language Models

Chuan Li, Qianyi Zhao, Fengran Mo, Cen Chen

专题命中 隐私与版权 :safety(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏

7. 安全评测 13 篇

2508.09937 2025-08-15 cs.CL cs.AI cs.LG 87%

A Comprehensive Evaluation framework of Alignment Techniques for LLMs

Muneeza Azmat, Momin Abbas, Maysa Malfiza Garcia de Macedo, Marcelo Carpinette Grave, Luan Soares de Souza, Tiago Machado, Rogerio A de Paula, Raya Horesh, Yixin Chen, Heloisa Caroline de Souza Pereira Candello, Rebecka Nordenlow, Aminat Adebiyi

机构 * IBM Research(IBM研究院)

专题命中 安全评测 :alignment(title,abstract);RLHF(abstract);safety(abstract);分类 cs.CL、cs.AI、cs.LG

Comments In submission

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.00399 2025-08-15 cs.CV 78%

iSafetyBench: A video-language benchmark for safety in industrial environment

Raiyaan Abdullah, Yogesh Singh Rawat, Shruti Vyas

机构 * University of Central Florida(中央佛罗里达大学)

专题命中 安全评测 :safety(title,abstract)

Comments Accepted to VISION'25 - ICCV 2025 workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10146 2025-08-15 cs.AI 70%

Agentic AI Frameworks: Architectures, Protocols, and Design Challenges

Hana Derouiche, Zaki Brahmi, Haithem Mazeni

机构 * University of Kairouan(卡鲁安大学) SMART Lab, University of Tunis(图纳大学SMART实验室) University of Sousse(索斯大学) Riadi Lab, Compus Manouba(曼努巴大学里亚迪实验室) University of Jandouba(贾杜巴大学)

专题命中 安全评测 :alignment(abstract);safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10028 2025-08-15 cs.CL cs.AI cs.HC cs.LG 67%

PREF: Reference-Free Evaluation of Personalised Text Generation in LLMs

Xiao Fu, Hossein A. Rahmani, Bin Wu, Jerome Ramos, Emine Yilmaz, Aldo Lipani

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments 7 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10494 2025-08-15 cs.LG cs.AI cs.MA 62%

A Unified Multi-Agent Framework for Universal Multimodal Understanding and Generation

Jiulin Li, Ping Huang, Yexin Li, Shuo Chen, Juewen Hu, Ye Tian

专题命中 安全评测 :alignment(abstract);分类 cs.AI、cs.LG

Comments 8 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10433 2025-08-15 cs.AI cs.CV cs.LG 62%

We-Math 2.0: A Versatile MathBook System for Incentivizing Visual Mathematical Reasoning

Runqi Qiao, Qiuna Tan, Peiqing Yang, Yanzi Wang, Xiaowan Wang, Enhui Wan, Sitong Zhou, Guanting Dong, Yuchen Zeng, Yida Xu, Jie Wang, Chong Sun, Chen Li, Honggang Zhang

机构 * BUPT(北京邮电大学) WeChat Vision, Tencent Inc.(腾讯公司) Tsinghua University(清华大学)

专题命中 安全评测 :alignment(abstract);分类 cs.AI、cs.LG

Comments Working in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10022 2025-08-15 cs.CL cs.AI 62%

Conformal P-Value in Multiple-Choice Question Answering Tasks with Provable Risk Control

Yuanchang Ye

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.13871 2025-08-15 cs.LG cs.AI cs.CR 62%

An Explainable Transformer-based Model for Phishing Email Detection: A Large Language Model Approach

Mohammad Amaz Uddin, Md Mahiuddin, Iqbal H. Sarker

机构 * Department of Computer Science and Engineering, BGC Trust University Bangladesh(Bangladesh BGC Trust 大学 计算机科学与工程系) Department of Computer Science and Engineering, International Islamic University Chittagong(Bangladesh 国际伊斯兰大学 昌德加荣分校 计算机科学与工程系) Centre for Securing Digital Futures, School of Science, Edith Cowan University(澳大利亚 埃德温·考文大学 科学学院 安全数字未来中心)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10869 2025-08-15 cs.CV cs.AI 57%

Medico 2025: Visual Question Answering for Gastrointestinal Imaging

Sushant Gautam, Vajira Thambawita, Michael Riegler, Pål Halvorsen, Steven Hicks

机构 * SimulaMet - Simula Metropolitan Center for Digital Engineering, Oslo, Norway(SimulaMet - Simula Metropolitan Center for Digital Engineering,挪威奥斯陆) Simula Research Laboratory, Oslo, Norway(Simula研究实验室,挪威奥斯陆) OsloMet - Oslo Metropolitan University, Oslo, Norway(OsloMet - 奥斯陆 Metropolitan 大学,挪威奥斯陆)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.13307 2025-08-15 cs.CV cs.AI 57%

Quantitative Comparison of Fine-Tuning Techniques for Pretrained Latent Diffusion Models in the Generation of Unseen SAR Images

Solène Debuysère, Nicolas Trouvé, Nathan Letheule, Olivier Lévêque, Elise Colin

机构 * Paris-Saclay University(巴黎-萨克雷大学) ONERA - The French Aerospace Lab(法国航空航天实验室)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10397 2025-08-15 cs.CV cs.AI 57%

PQ-DAF: Pose-driven Quality-controlled Data Augmentation for Data-scarce Driver Distraction Detection

Haibin Sun, Xinghui Song

机构 * College of Computer Science and Engineering, Shandong University of Science and Technology(计算机科学与工程学院,山东科技大学)

专题命中 安全评测 :safety(abstract);分类 cs.AI

Comments 11 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10358 2025-08-15 cs.AI 57%

What to Ask Next? Probing the Imaginative Reasoning of LLMs with TurtleSoup Puzzles

Mengtao Zhou, Sifan Wu, Huan Zhang, Qi Sima, Bang Liu

专题命中 安全评测 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10309 2025-08-15 cs.CV 50%

From Pixel to Mask: A Survey of Out-of-Distribution Segmentation

Wenjie Zhao, Jia Li, Yunhui Guo

机构 * University of Texas at Dallas(德克萨斯大学达拉斯分校)

专题命中 安全评测 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

8. AI治理与伦理 1 篇

2508.10007 2025-08-15 cs.CL stat.ME 57%

Automated scoring of the Ambiguous Intentions Hostility Questionnaire using fine-tuned large language models

Y. Lyu, D. Combs, D. Neumann, Y. C. Leong

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL

Comments We have no known conflict of interest

详情

展开后加载摘要…

URL PDF HTML 收藏