arXivDaily arXiv每日学术速递 周一至周五更新
今日文章,稍等片刻,小憩一会儿再来看看

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-08-07 至 2025-08-07 共收录 16 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 16 篇

2508.04251 2025-08-07 cs.LG 79%

T3Time: Tri-Modal Time Series Forecasting via Adaptive Multi-Head Alignment and Residual Fusion

Abdul Monaf Chowdhury, Rabeya Akter, Safaeid Hossain Arib

专题命中 安全评测 :alignment(title,abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04240 2025-08-07 eess.SP 78%

ChineseEEG-2: An EEG Dataset for Multimodal Semantic Alignment and Neural Decoding during Reading and Listening

Sitong Chen, Beiqianyi Li, Cuilin He, Dongyang Li, Mingyang Wu, Xinke Shen, Song Wang, Xuetao Wei, Xindi Wang, Haiyan Wu, Quanying Liu

专题命中 安全评测 :alignment(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03936 2025-08-07 cs.CR cs.CL cs.LG cs.SE 73%

ASTRA: Autonomous Spatial-Temporal Red-teaming for AI Software Assistants

Xiangzhe Xu, Guangyu Shen, Zian Su, Siyuan Cheng, Hanxi Guo, Lu Yan, Xuan Chen, Jiasheng Jiang, Xiaolong Jin, Chengpeng Wang, Zhuo Zhang, Xiangyu Zhang

专题命中 安全评测 :alignment(abstract);safety(abstract);分类 cs.CL、cs.LG

Comments The first two authors (Xiangzhe Xu and Guangyu Shen) contributed equally to this work

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04350 2025-08-07 cs.CL cs.AI cs.CV cs.LG cs.MA 67%

Chain of Questions: Guiding Multimodal Curiosity in Language Models

Nima Iji, Kia Dashtipour

机构 * Edinburgh Napier University(爱丁堡纳皮尔大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04442 2025-08-07 cs.CL cs.AI 62%

Automated Generation of Curriculum-Aligned Multiple-Choice Questions for Malaysian Secondary Mathematics Using Generative AI

Rohaizah Abdul Wahid, Muhamad Said Nizamuddin Nadim, Suliana Sulaiman, Syahmi Akmal Shaharudin, Muhammad Danial Jupikil, Iqqwan Jasman Su Azlan Su

机构 * Fakulti Komputeran dan Meta-Teknologi (META), Universiti Pendidikan Sultan Idris(计算机与元技术学院(META),苏丹依德里斯教育大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03714 2025-08-07 cs.HC cs.AI cs.CR cs.CY 62%

"Think First, Verify Always": Training Humans to Face AI Risks

Yuksel Aydin

机构 * Independent Researcher(独立研究者)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04469 2025-08-07 cs.CV cs.CL 57%

FrEVL: Leveraging Frozen Pretrained Embeddings for Efficient Vision-Language Understanding

Emmanuelle Bourigault, Pauline Bourigault

机构 * Department of Engineering Science, University of Oxford(牛津大学工程科学系) Department of Electrical Engineering, Imperial College London(伦敦帝国学院电子工程系)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

Comments 8 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04199 2025-08-07 cs.CL 57%

Reasoning Beyond Labels: Measuring LLM Sentiment in Low-Resource, Culturally Nuanced Contexts

Millicent Ochieng, Anja Thieme, Ignatius Ezeani, Risa Ueno, Samuel Maina, Keshet Ronen, Javier Gonzalez, Jacki O'Neill

机构 * Microsoft Research(微软研究院) Lancaster University(兰卡斯特大学) University of Washington(华盛顿大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04011 2025-08-07 cs.HC cs.AI 57%

StepWrite: Adaptive Planning for Speech-Driven Text Generation

Hamza El Alaoui, Atieh Taheri, Yi-Hao Peng, Jeffrey P. Bigham

机构 * School of Computer Science Carnegie Mellon University Pittsburgh PA USA School of Computer Science Carnegie Mellon University

专题命中 安全评测 :alignment(abstract);分类 cs.AI

Comments This paper has been accepted to UIST 2025. For additional materials and project details, please see: https://www.cs.cmu.edu/~helalaou/publications/stepwrite

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03734 2025-08-07 eess.IV cs.AI cs.CV 57%

A Survey of Multimodal Ophthalmic Diagnostics: From Task-Specific Approaches to Foundational Models

Xiaoling Luo, Ruli Zheng, Qiaojian Zheng, Zibo Du, Shuo Yang, Meidan Ding, Qihao Xu, Chengliang Liu, Linlin Shen

机构 * College of Computer Science and Software Engineering, Shenzhen University, Shenzhen, China(深圳大学计算机科学与软件工程学院) Shenzhen Key Laboratory of Visual Object Detection and Recognition, Harbin Institute of Technology, Shenzhen, 518055, China(视觉对象检测与识别深圳重点实验室) Laboratory for Artificial Intelligence in Design, Hong Kong(人工智能设计实验室) School of Artificial Intelligence, Shenzhen University, Shenzhen, China(深圳大学人工智能学院)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03722 2025-08-07 cs.CV cs.AI 57%

Multimodal Video Emotion Recognition with Reliable Reasoning Priors

Zhepeng Wang, Yingjian Zhu, Guanghao Dong, Hongzhu Yi, Feng Chen, Xinming Wang, Jun Xie

机构 * Lenovo Research(联想研究院) School of Artificial Intelligence, UCAS(中国科学院大学人工智能学院) Institute of Automation, CAS(中国科学院自动化研究所) Macau University of Science and Technology(澳门科学理工学院) School of Computer Science and Technology, UCAS(中国科学院大学计算机科学与技术学院)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

Comments preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.13783 2025-08-07 cs.MA cs.AI cs.GT cs.SY eess.SY 57%

A Value Based Parallel Update MCTS Method for Multi-Agent Cooperative Decision Making of Connected and Automated Vehicles

Ye Han, Lijun Zhang, Dejian Meng, Zhuang Zhang, Xingyu Hu, Songyu Weng

机构 * School of Automotive Studies, Tongji University(同济大学汽车学院)

专题命中 安全评测 :safety(abstract);分类 cs.AI

Comments arXiv admin note: text overlap with arXiv:2408.04295 by other authors

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.06083 2025-08-07 cs.LG cs.IR 57%

A Survey of Controllable Learning: Methods and Applications in Information Retrieval

Chenglei Shen, Xiao Zhang, Teng Shi, Changshuo Zhang, Guofu Xie, Jun Xu

机构 * Gaoling School of Artificial Intelligence, Renmin University of China, Beijing 100872, China(中国人民大学人工智能学院) AI Lab at Lenovo Research, Beijing 100085, China(联想研究院人工智能实验室)

专题命中 安全评测 :trustworthy(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04450 2025-08-07 eess.IV cs.CV 50%

TotalRegistrator: Towards a Lightweight Foundation Model for CT Image Registration

Xuan Loc Pham, Gwendolyn Vuurberg, Marjan Doppen, Joey Roosen, Tip Stille, Thi Quynh Ha, Thuy Duong Quach, Quoc Vu Dang, Manh Ha Luu, Ewoud J. Smit, Hong Son Mai, Mattias Heinrich, Bram van Ginneken, Mathias Prokop, Alessa Hering

机构 * Department of Imaging, Radboudumc, Nijmegen, the Netherlands(影像部,Radboudumc,尼姆egen,荷兰) Institute for Medical Informatics, University of Lübeck, Lübeck, Germany(医学信息研究所,吕贝克大学,吕贝克,德国) Department of Nuclear Medicine, Hospital 108, Hanoi, Vietnam(核医学部,108医院,河内,越南) Department of Diagnostic Imaging and Interventional Radiology, Thai Nguyen National Hospital, Thai Nguyen, Vietnam(诊断影像和介入放射科,Thai Nguyen国家医院,Thai Nguyen,越南) Interventional Radiology Center, Tam Anh Hospital, Hanoi, Vietnam(介入放射科中心,Tam Anh医院,河内,越南) FET, Vietnam National University, University of Engineering and Technology, Hanoi, Vietnam(FET,越南国家大学,工程与技术大学,河内,越南)

专题命中 安全评测 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04448 2025-08-07 cs.SE 50%

Large Language Models Versus Static Code Analysis Tools: A Systematic Benchmark for Vulnerability Detection

Damian Gnieciak, Tomasz Szandala

专题命中 安全评测 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04120 2025-08-07 cs.CV 50%

CLIPVehicle: A Unified Framework for Vision-based Vehicle Search

Likai Wang, Ruize Han, Xiangqun Zhang, Wei Feng

机构 * College of Intelligence and Computing, Tianjin University(智能与计算学院,天津大学) Shenzhen Institute of Advanced Technology, Chinese Academy of Sciences(深圳先进技术研究院,中国科学院) Shenzhen University of Advanced Technology(深圳大学先进技术学院)

专题命中 安全评测 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏