arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-08-19 至 2025-08-19 共收录 20 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 20 篇

2505.11247 2025-08-19 cs.AI cs.LG cs.RO 81%

LD-Scene: LLM-Guided Diffusion for Controllable Generation of Adversarial Safety-Critical Driving Scenarios

Mingxing Peng, Yuting Xie, Xusen Guo, Ruoyu Yao, Hai Yang, Jun Ma

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) School of Computer Science and Engineering, Sun Yat-sen University(中山大学计算机科学与工程学院)

专题命中 安全评测 :safety(title,abstract);分类 cs.AI、cs.LG

Comments 18 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.12624 2025-08-19 cs.CL cs.AI 81%

Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges

Aman Singh Thakur, Kartik Choudhary, Venkat Srinik Ramayapally, Sankaran Vaidyanathan, Dieuwke Hupkes

机构 * University of Massachusetts Amherst(马萨诸塞大学阿默斯特分校) Meta

专题命中 安全评测 :alignment(title,abstract);分类 cs.CL、cs.AI

Comments https://aclanthology.org/2025.gem-1.33/

Journal ref Proceedings of the Fourth Workshop on Generation Evaluation and Metrics GEM2 2025 pages 404 to 430; July 31 August 1 2025; 2025 Association for Computational Linguistics

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.10872 2025-08-19 cs.CV cs.AI cs.ET 79%

V-RoAst: Visual Road Assessment. Can VLM be a Road Safety Assessor Using the iRAP Standard?

Natchapon Jongwiriyanurak, Zichao Zeng, June Moh Goo, Xinglei Wang, Ilya Ilyankou, Kerkritt Sriroongvikrai, Nicola Christie, Meihui Wang, Huanfa Chen, James Haworth

机构 * University College London(伦敦大学学院) Chulalongkorn University(朱拉隆梭大学)

专题命中 安全评测 :safety(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.01408 2025-08-19 cs.RO cs.CV 78%

From Shadows to Safety: Occlusion Tracking and Risk Mitigation for Urban Autonomous Driving

Korbinian Moller, Luis Schwarzmeier, Johannes Betz

专题命中 安全评测 :safety(title,abstract)

Comments 8 Pages. Submitted to the IEEE Intelligent Vehicles Symposium (IV 2025), Romania

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.16867 2025-08-19 cs.CV 78%

ETVA: Evaluation of Text-to-Video Alignment via Fine-grained Question Generation and Answering

Kaisi Guan, Zhengfeng Lai, Yuchong Sun, Peng Zhang, Wei Liu, Kieran Liu, Meng Cao, Ruihua Song

机构 * Renmin University of China(中国人民大学) Apple(苹果公司)

专题命中 安全评测 :alignment(title,abstract)

Comments International Conference on Computer Vision 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11957 2025-08-19 cs.MA cs.AI cs.LG 73%

A Comprehensive Review of AI Agents: Transforming Possibilities in Technology and Beyond

Xiaodong Qu, Andrews Damoah, Joshua Sherwood, Peiyan Liu, Christian Shun Jin, Lulu Chen, Minjie Shen, Nawwaf Aleisa, Zeyuan Hou, Chenyu Zhang, Lifu Gao, Yanshu Li, Qikai Yang, Qun Wang, Cristabelle De Souza

机构 * University of Maryland, College Park(马里兰大学学院公园分校) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Brown University(布朗大学) San Francisco State University(旧金山州立大学) Stanford University(斯坦福大学)

专题命中 安全评测 :safety(abstract);trustworthy(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06225 2025-08-19 cs.AI 70%

Overconfidence in LLM-as-a-Judge: Diagnosis and Confidence-Driven Solution

Zailong Tian, Zhuoheng Han, Yanzhe Chen, Haozhe Xu, Xi Yang, Richeng Xuan, Houfeng Wang, Lizi Liao

专题命中 安全评测 :alignment(abstract);trustworthy(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11647 2025-08-19 cs.LO cs.AI 70%

Categorical Construction of Logically Verifiable Neural Architectures

Logan Nye

机构 * Carnegie Mellon University School of Computer Science(卡内基梅隆大学计算机科学学院)

专题命中 安全评测 :safety(abstract);trustworthy(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.13152 2025-08-19 cs.CL cs.AI 62%

RepreGuard: Detecting LLM-Generated Text by Revealing Hidden Representation Patterns

Xin Chen, Junchao Wu, Shu Yang, Runzhe Zhan, Zeyu Wu, Ziyang Luo, Di Wang, Min Yang, Lidia S. Chao, Derek F. Wong

机构 * NLP(自然语言处理) CT Lab, Department of Computer and Information Science, University of Macau(计算机与信息科学系,澳门大学) Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences(深圳先进技术研究院,中国科学院) Provable Responsible AI and Data Analytics Lab, KAUST(可证明责任AI与数据分析实验室,卡尔斯兰大学) Hong Kong Baptist University(香港 Baptist 大学)

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL、cs.AI

Comments Accepted to TACL 2025. This version is a pre-MIT Press publication version

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12212 2025-08-19 cs.LG cs.AI q-bio.QM 62%

ProtTeX-CC: Activating In-Context Learning in Protein LLM via Two-Stage Instruction Compression

Chuanliu Fan, Zicheng Ma, Jun Gao, Nan Yu, Jun Zhang, Ziqiang Cao, Yi Qin Gao, Guohong Fu

专题命中 安全评测 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.04596 2025-08-19 cs.AI cs.CE cs.CL 62%

SECQUE: A Benchmark for Evaluating Real-World Financial Analysis Capabilities

Noga Ben Yoash, Meni Brief, Oded Ovadia, Gil Shenderovitz, Moshik Mishaeli, Rachel Lemberg, Eitam Sheetrit

机构 * Microsoft Industry AI(微软产业人工智能)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

Comments Benchmark available at: https://huggingface.co/datasets/nogabenyoash/SecQue

Journal ref n Proceedings of the Fourth Workshop on Generation, Evaluation and Metrics, Association for Computational Linguistics (2025) https://aclanthology.org/2025.gem-1.16/

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12282 2025-08-19 cs.CL cs.IR 57%

A Question Answering Dataset for Temporal-Sensitive Retrieval-Augmented Generation

Ziyang Chen, Erxue Min, Xiang Zhao, Yunxin Li, Xin Jia, Jinzhi Liao, Jichao Li, Shuaiqiang Wang, Baotian Hu, Dawei Yin

机构 * Laboratory for Big Data and Decision, National University of Defense Technology, Changsha, China(大数据与决策实验室,国防科技大学,长沙,中国) Baidu Inc., Beijing, China(百度公司,北京,中国) Department of Computer Science and Technology, Harbin Institute of Technology (Shenzhen), Shenzhen, China(计算机科学与技术系,哈尔滨工业大学(深圳),深圳,中国)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

Comments 10 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12100 2025-08-19 cs.AI 57%

Overcoming Knowledge Discrepancies: Structuring Reasoning Threads through Knowledge Balancing in Interactive Scenarios

Daniel Burkhardt, Xiangwei Cheng

机构 * Ferdinand Steinbeis Institute(费尔迪南·斯坦贝茨研究所)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

Comments 13 pages, 1 figure, 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.15125 2025-08-19 cs.AI 57%

Contemplative Artificial Intelligence

Ruben Laukkonen, Fionn Inglis, Shamil Chandaria, Lars Sandved-Smith, Edmundo Lopez-Sola, Jakob Hohwy, Jonathan Gold, Adam Elwood

专题命中 安全评测 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11848 2025-08-19 quant-ph cs.ET cs.LG 57%

Adversarial Robustness in Distributed Quantum Machine Learning

Pouya Kananian, Hans-Arno Jacobsen

机构 * Department of Electrical and Computer Engineering, University of Toronto(电气与计算机工程系,多伦多大学)

专题命中 安全评测 :trustworthy(abstract);分类 cs.LG

Comments This is a preprint of a book chapter that is planned to be published in "Quantum Robustness in Artificial Intelligence" by Springer Nature

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.03603 2025-08-19 cs.LG 57%

MUC: Machine Unlearning for Contrastive Learning with Black-box Evaluation

Yihan Wang, Yiwei Lu, Guojun Zhang, Franziska Boenisch, Adam Dziedzic, Yaoliang Yu, Xiao-Shan Gao

机构 * University of Waterloo(滑铁卢大学) University of Ottawa(渥太华大学) Alibaba(阿里巴巴) CISPA Helmholtz Center for Information Security(信息安全赫尔姆霍兹中心) Vector Institute(向量研究所) Academy of Mathematics and Systems Science, Chinese Academy of Sciences(中国科学院数学与系统科学研究院) University of Chinese Academy of Sciences(中国科学院大学)

专题命中 安全评测 :alignment(abstract);分类 cs.LG

Comments Published in TMLR

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.13954 2025-08-19 cs.CL 57%

Measuring Social Biases in Masked Language Models by Proxy of Prediction Quality

Rahul Zalkikar, Kanchan Chandra

机构 * New York University(纽约大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

Journal ref Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2025, pages 1337--1361

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12752 2025-08-19 cs.IR 50%

Deep Research: A Survey of Autonomous Research Agents

Wenlin Zhang, Xiaopeng Li, Yingyi Zhang, Pengyue Jia, Yichao Wang, Huifeng Guo, Yong Liu, Xiangyu Zhao

专题命中 安全评测 :trustworthy(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12384 2025-08-19 cs.CV cs.CR 50%

ViT-EnsembleAttack: Augmenting Ensemble Models for Stronger Adversarial Transferability in Vision Transformers

Hanwen Cao, Haobo Lu, Xiaosen Wang, Kun He

机构 * School of Computer Science and Technology(计算机科学与技术学院) Huazhong University of Science and Technology(华中科技大学)

专题命中 安全评测 :trustworthy(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21610 2025-08-19 cs.RO cs.CV 50%

Research Challenges and Progress in the End-to-End V2X Cooperative Autonomous Driving Competition

Ruiyang Hao, Haibao Yu, Jiaru Zhong, Chuanye Wang, Jiahao Wang, Yiming Kan, Wenxian Yang, Siqi Fan, Huilin Yin, Jianing Qiu, Yao Mu, Jiankai Sun, Li Chen, Walter Zimmer, Dandan Zhang, Shanghang Zhang, Mac Schwager, Ping Luo, Zaiqing Nie

机构 * Tsinghua University(清华大学) Hong Kong University(香港大学) Tongji University(同济大学) Chinese University of Hong Kong(香港中文大学) Shanghai Jiao Tong University(上海交通大学) Stanford University(斯坦福大学) OpenDriveLab(OpenDrive实验室) Technical University of Munich(慕尼黑技术大学) Imperial College London(伦敦帝国理工学院) Peking University(北京大学)

专题命中 安全评测 :safety(abstract)

Comments 10 pages, 4 figures, accepted by ICCVW Author list updated to match the camera-ready version, in compliance with conference policy

详情

展开后加载摘要…

URL PDF HTML 收藏