arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-11-19 至 2025-11-19 共收录 9 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 其他安全 9 篇

2511.13909 2025-11-19 cs.CV 78%

Mind the Gap: Evaluating LLM Understanding of Human-Taught Road Safety Principles

Chalamalasetti Kranti

机构 * uni-potsdam(波恩大学)

专题命中 其他安全 :safety(title,abstract)

Comments Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.06009 2025-11-19 cs.CV 78%

Continual Learning for Image Captioning through Improved Image-Text Alignment

Bertram Taetz, Gal Bordelius

机构 * IT & Engineering International University of Applied Sciences(IT与工程国际应用科学大学)

专题命中 其他安全 :alignment(title,abstract)

Comments 11 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13761 2025-11-19 cs.DC cs.AI cs.LG 62%

What happens when nanochat meets DiLoCo?

Alexander Acker, Soeren Becker, Sasho Nedelkoski, Dominik Scheinert, Odej Kao, Philipp Wiesner

机构 * Team exalsius(exalsius团队) Technische Universität Berlin(柏林技术大学)

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

Comments 8pages, 3 figures, technical report

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01293 2025-11-19 cs.CV cs.AI 57%

GMAT: Grounded Multi-Agent Clinical Description Generation for Text Encoder in Vision-Language MIL for Whole Slide Image Classification

Ngoc Bui Lam Quang, Nam Le Nguyen Binh, Thanh-Huy Nguyen, Le Thien Phuc Nguyen, Quan Nguyen, Ulas Bagci

机构 * AI VIETNAM(AI越南) Carnegie Mellon University(卡内基梅隆大学) University of Wisconsin-Madison(威斯康星大学麦迪逊分校) PTIT Northwestern University(西北大学)

专题命中 其他安全 :alignment(abstract);分类 cs.AI

Comments Acccepted in MICCAI Workshop 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.01916 2025-11-19 cs.RO cs.LG 57%

Generalizable and Fast Surrogates: Model Predictive Control of Articulated Soft Robots using Physics-Informed Neural Networks

Tim-Lukas Habich, Aran Mohammad, Simon F. G. Ehlers, Martin Bensch, Thomas Seel, Moritz Schappler

机构 * Leibniz University Hannover(莱布尼茨汉诺威大学)

专题命中 其他安全 :safety(abstract);分类 cs.LG

Comments Accepted for publication in IEEE Transactions on Robotics (T-RO) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12511 2025-11-19 cs.CV cs.LG 57%

DINO-Detect: A Simple yet Effective Framework for Blur-Robust AI-Generated Image Detection

Jialiang Shen, Jiyang Zheng, Yunqi Xue, Huajie Chen, Yu Yao, Hui Kang, Ruiqi Liu, Helin Gong, Yang Yang, Dadong Wang, Tongliang Liu

机构 * Sydney AI Center, The University of Sydney(悉尼人工智能中心,悉尼大学) CSIRO, Data61(澳大利亚联邦科学与工业研究组织、Data61) Shanghai Jiao Tong University(上海交通大学) City University of Macau(澳门城市大学) CASIA(中国科学院自动化研究所)

专题命中 其他安全 :safety(abstract);分类 cs.LG

Comments 12 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.07991 2025-11-19 cs.AI 57%

VSPO: Validating Semantic Pitfalls in Ontology via LLM-Based CQ Generation

Hyojun Choi, Seokju Hwang, Kyong-Ho Lee

专题命中 其他安全 :alignment(abstract);分类 cs.AI

Comments Accepted at AAAI 2026 oral

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13918 2025-11-19 cs.HC 50%

Human-centric Maintenance Process Through Integration of AI, Speech, and AR

Parul Khanna, Ravdeep Kour, Ramin Karim

专题命中 其他安全 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13876 2025-11-19 cs.CV 50%

QwenCLIP: Boosting Medical Vision-Language Pretraining via LLM Embeddings and Prompt tuning

Xiaoyang Wei, Camille Kurtz, Florence Cloppet

机构 * Laboratoire d'Informatique Paris Descartes (LIPADE), Université Paris Cité (France)(巴黎笛卡尔大学信息学实验室(LIPADE),巴黎城市大学(法国))

专题命中 其他安全 :alignment(abstract)

Comments This work has been submitted to the IEEE ISBI for possible publication

详情

展开后加载摘要…

URL PDF HTML 收藏