arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-10-29 至 2025-10-29 共收录 16 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 16 篇

2505.11842 2025-10-29 cs.CV cs.CL 79%

Video-SafetyBench: A Benchmark for Safety Evaluation of Video LVLMs

Xuannan Liu, Zekun Li, Zheqi He, Peipei Li, Shuhan Xia, Xing Cui, Huaibo Huang, Xi Yang, Ran He

机构 * Beijing University of Posts and Telecommunications(北京邮电大学) Beijing Academy of Artificial Intelligence(北京人工智能研究院) University of California, Santa Barbara(加州大学圣芭芭拉分校) Center for Research on Intelligent Perception and Computing, NLPR, CASIA(中国科学院CASIA智能感知与计算中心)

专题命中 安全评测 :safety(title,abstract);分类 cs.CL

Comments Accepted by NeurIPS 2025 Dataset and Benchmark Track, Project page: https://liuxuannan.github.io/Video-SafetyBench.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.01128 2025-10-29 cs.CV cs.AI 74%

RipVIS: Rip Currents Video Instance Segmentation Benchmark for Beach Monitoring and Safety

Andrei Dumitriu, Florin Tatui, Florin Miron, Aakash Ralhan, Radu Tudor Ionescu, Radu Timofte

机构 * Computer Vision Lab, CAIDAS & IFI, University of Würzburg, Germany(计算机视觉实验室,CAIDAS与IFI,乌尔姆大学,德国) University of Bucharest, Romania(布加勒斯特大学,罗马尼亚)

专题命中 安全评测 :safety(title);分类 cs.AI

Comments Accepted at CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23921 2025-10-29 cs.CL cs.LG 73%

Breaking the Benchmark: Revealing LLM Bias via Minimal Contextual Augmentation

Kaveh Eskandari Miandoab, Mahammed Kamruzzaman, Arshia Gharooni, Gene Louis Kim, Vasanth Sarathy, Ninareh Mehrabi

机构 * Tufts University(塔夫茨大学) University of South Florida(佛罗里达州立大学) Sharif University of Technology(谢赫·穆吉布技术大学) Meta Resolution(解决方案)

专题命中 安全评测 :alignment(abstract);safety(abstract);分类 cs.CL、cs.LG

Comments 9 pages, 3 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23966 2025-10-29 cs.LG cs.SE 70%

A Pragmatic Way to Measure Chain-of-Thought Monitorability

Scott Emmons, Roland S. Zimmermann, David K. Elson, Rohin Shah

机构 * Google(谷歌)

专题命中 安全评测 :safety(abstract);AI safety(abstract);分类 cs.LG

Comments The first two authors contributed equally

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.24488 2025-10-29 cs.CL cs.AI 62%

A word association network methodology for evaluating implicit biases in LLMs compared to humans

Katherine Abramski, Giulio Rossetti, Massimo Stella

机构 * University of Pisa, Department of Computer Science(比萨大学计算机科学系) National Research Council of Italy, Institute of Information Science and Technologies(意大利国家研究委员会信息科学与技术研究所) University of Trento, Department of Psychology and Cognitive Science(特伦托大学心理学与认知科学系)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

Comments 24 pages, 13 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.13499 2025-10-29 cs.CY cs.AI 62%

Reproducible workflow for online AI in digital health

Susobhan Ghosh, Bhanu T. Gullapalli, Daiqi Gao, Asim Gazi, Anna Trella, Ziping Xu, Kelly Zhang, Susan A. Murphy

机构 * Harvard University(哈佛大学) Imperial College London(帝国理工学院)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.24115 2025-10-29 cs.AI cs.LG 62%

HistoLens: An Interactive XAI Toolkit for Verifying and Mitigating Flaws in Vision-Language Models for Histopathology

Sandeep Vissapragada, Vikrant Sahu, Gagan Raj Gupta, Vandita Singh

机构 * Indian Institute of Technology(印度理工学院) All India Institute of Medical Sciences(全印度医学科学研究所)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23854 2025-10-29 cs.CL cs.AI 62%

Can LLMs Narrate Tabular Data? An Evaluation Framework for Natural Language Representations of Text-to-SQL System Outputs

Jyotika Singh, Weiyi Sun, Amit Agarwal, Viji Krishnamurthy, Yassine Benajiba, Sujith Ravi, Dan Roth

机构 * Oracle AI

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

Comments Accepted at EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.24431 2025-10-29 cs.IR cs.AI 57%

MiniOneRec: An Open-Source Framework for Scaling Generative Recommendation

Xiaoyu Kong, Leheng Sheng, Junfei Tan, Yuxin Chen, Jiancan Wu, An Zhang, Xiang Wang, Xiangnan He

机构 * University of Science and Technology of China(中国科学技术大学) National University of Singapore(新加坡国立大学)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

Comments Technical Report

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20927 2025-10-29 cs.CY 57%

What do model reports say about their ChemBio benchmark evaluations? Comparing recent releases to the STREAM framework

Tom Reed, Tegan McCaslin, Luca Righetti

专题命中 安全评测 :safety(abstract);分类 cs.CY

Comments 12 pages, 6 figures. Includes appendices Added supplementary materials

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.01778 2025-10-29 cs.LG stat.ML 57%

CONFIDERAI: a novel CONFormal Interpretable-by-Design score function for Explainable and Reliable Artificial Intelligence

Sara Narteni, Alberto Carlevaro, Fabrizio Dabbene, Marco Muselli, Maurizio Mongelli

机构 * Institute of Electronics, Information Engineering and Telecommunications - National Research Council of Italy (CNR-IEIIT)(电子信息与电信研究所 - 意大利国家科研理事会(CNR-IEIIT)) Politecnico di Torino - Department of Control and Computer Engineering (DAUIN)(托尼诺理工学院 - 控制与计算机工程系(DAUIN)) Rulex Innovation Labs(Rulex创新实验室)

专题命中 安全评测 :trustworthy(abstract);分类 cs.LG

Comments 26 pages, 3 figures, international journal

Journal ref Pattern Recognition, Volume 171, Part B, 2026, 112219, ISSN 0031-3203

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.24259 2025-10-29 cs.CL cs.RO 57%

Can LLMs Translate Human Instructions into a Reinforcement Learning Agent's Internal Emergent Symbolic Representation?

Ziqi Ma, Sao Mai Nguyen, Philippe Xu

机构 * U2IS, ENSTA, IP-Paris(U2IS、ENSTA、IP-巴黎)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23622 2025-10-29 cs.LG cs.CR 57%

Adversarially-Aware Architecture Design for Robust Medical AI Systems

Alyssa Gerhart, Balaji Iyangar

机构 * Department of Computer Science Benedict College Columbia, SC, USA(计算机科学系 奥本大学 奥斯本, 南卡罗来纳州, 美国)

专题命中 安全评测 :safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23988 2025-10-29 cs.RO 50%

A Survey on Collaborative SLAM with 3D Gaussian Splatting

Phuc Nguyen Xuan, Thanh Nguyen Canh, Huu-Hung Nguyen, Nak Young Chong, Xiem HoangVan

机构 * Institute of System Integration, Le Quy Don Technical University, Hanoi, 10000, Vietnam(越南吕文技术大学系统整合研究所) Graduate School of Advanced Science and Technology, Japan Advanced Institute of Science and Technology(日本先进科学技术研究院先进科学技术研究生院) University of Engineering and Technology, Vietnam National University, Hanoi, 10000, Vietnam(越南国家大学工程技术大学)

专题命中 安全评测 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23908 2025-10-29 eess.SP cs.ET 50%

Machine Learning-Driven User Localization in RIS-Assisted Wireless Systems

M. T. Hassan, D. Zelenchuk, M. A. B. Abbasi

专题命中 安全评测 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.12841 2025-10-29 cs.CV 50%

AnyCap Project: A Unified Framework, Dataset, and Benchmark for Controllable Omni-modal Captioning

Yiming Ren, Zhiqiang Lin, Yu Li, Gao Meng, Weiyun Wang, Junjie Wang, Zicheng Lin, Jifeng Dai, Yujiu Yang, Wenhai Wang, Ruihang Chu

机构 * Tsinghua University(清华大学) Shanghai AI Laboratory(上海人工智能实验室) Fudan University(复旦大学) The Chinese University of Hong Kong(香港中文大学)

专题命中 安全评测 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏