arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-08-19 至 2025-08-19 共收录 58 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 20 篇

2504.04596 2025-08-19 cs.AI cs.CE cs.CL 62%

SECQUE: A Benchmark for Evaluating Real-World Financial Analysis Capabilities

Noga Ben Yoash, Meni Brief, Oded Ovadia, Gil Shenderovitz, Moshik Mishaeli, Rachel Lemberg, Eitam Sheetrit

机构 * Microsoft Industry AI(微软产业人工智能)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

Comments Benchmark available at: https://huggingface.co/datasets/nogabenyoash/SecQue

Journal ref n Proceedings of the Fourth Workshop on Generation, Evaluation and Metrics, Association for Computational Linguistics (2025) https://aclanthology.org/2025.gem-1.16/

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12282 2025-08-19 cs.CL cs.IR 57%

A Question Answering Dataset for Temporal-Sensitive Retrieval-Augmented Generation

Ziyang Chen, Erxue Min, Xiang Zhao, Yunxin Li, Xin Jia, Jinzhi Liao, Jichao Li, Shuaiqiang Wang, Baotian Hu, Dawei Yin

机构 * Laboratory for Big Data and Decision, National University of Defense Technology, Changsha, China(大数据与决策实验室,国防科技大学,长沙,中国) Baidu Inc., Beijing, China(百度公司,北京,中国) Department of Computer Science and Technology, Harbin Institute of Technology (Shenzhen), Shenzhen, China(计算机科学与技术系,哈尔滨工业大学(深圳),深圳,中国)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

Comments 10 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12100 2025-08-19 cs.AI 57%

Overcoming Knowledge Discrepancies: Structuring Reasoning Threads through Knowledge Balancing in Interactive Scenarios

Daniel Burkhardt, Xiangwei Cheng

机构 * Ferdinand Steinbeis Institute(费尔迪南·斯坦贝茨研究所)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

Comments 13 pages, 1 figure, 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.15125 2025-08-19 cs.AI 57%

Contemplative Artificial Intelligence

Ruben Laukkonen, Fionn Inglis, Shamil Chandaria, Lars Sandved-Smith, Edmundo Lopez-Sola, Jakob Hohwy, Jonathan Gold, Adam Elwood

专题命中 安全评测 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11848 2025-08-19 quant-ph cs.ET cs.LG 57%

Adversarial Robustness in Distributed Quantum Machine Learning

Pouya Kananian, Hans-Arno Jacobsen

机构 * Department of Electrical and Computer Engineering, University of Toronto(电气与计算机工程系,多伦多大学)

专题命中 安全评测 :trustworthy(abstract);分类 cs.LG

Comments This is a preprint of a book chapter that is planned to be published in "Quantum Robustness in Artificial Intelligence" by Springer Nature

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.03603 2025-08-19 cs.LG 57%

MUC: Machine Unlearning for Contrastive Learning with Black-box Evaluation

Yihan Wang, Yiwei Lu, Guojun Zhang, Franziska Boenisch, Adam Dziedzic, Yaoliang Yu, Xiao-Shan Gao

机构 * University of Waterloo(滑铁卢大学) University of Ottawa(渥太华大学) Alibaba(阿里巴巴) CISPA Helmholtz Center for Information Security(信息安全赫尔姆霍兹中心) Vector Institute(向量研究所) Academy of Mathematics and Systems Science, Chinese Academy of Sciences(中国科学院数学与系统科学研究院) University of Chinese Academy of Sciences(中国科学院大学)

专题命中 安全评测 :alignment(abstract);分类 cs.LG

Comments Published in TMLR

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.13954 2025-08-19 cs.CL 57%

Measuring Social Biases in Masked Language Models by Proxy of Prediction Quality

Rahul Zalkikar, Kanchan Chandra

机构 * New York University(纽约大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

Journal ref Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2025, pages 1337--1361

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12752 2025-08-19 cs.IR 50%

Deep Research: A Survey of Autonomous Research Agents

Wenlin Zhang, Xiaopeng Li, Yingyi Zhang, Pengyue Jia, Yichao Wang, Huifeng Guo, Yong Liu, Xiangyu Zhao

专题命中 安全评测 :trustworthy(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12384 2025-08-19 cs.CV cs.CR 50%

ViT-EnsembleAttack: Augmenting Ensemble Models for Stronger Adversarial Transferability in Vision Transformers

Hanwen Cao, Haobo Lu, Xiaosen Wang, Kun He

机构 * School of Computer Science and Technology(计算机科学与技术学院) Huazhong University of Science and Technology(华中科技大学)

专题命中 安全评测 :trustworthy(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21610 2025-08-19 cs.RO cs.CV 50%

Research Challenges and Progress in the End-to-End V2X Cooperative Autonomous Driving Competition

Ruiyang Hao, Haibao Yu, Jiaru Zhong, Chuanye Wang, Jiahao Wang, Yiming Kan, Wenxian Yang, Siqi Fan, Huilin Yin, Jianing Qiu, Yao Mu, Jiankai Sun, Li Chen, Walter Zimmer, Dandan Zhang, Shanghang Zhang, Mac Schwager, Ping Luo, Zaiqing Nie

机构 * Tsinghua University(清华大学) Hong Kong University(香港大学) Tongji University(同济大学) Chinese University of Hong Kong(香港中文大学) Shanghai Jiao Tong University(上海交通大学) Stanford University(斯坦福大学) OpenDriveLab(OpenDrive实验室) Technical University of Munich(慕尼黑技术大学) Imperial College London(伦敦帝国理工学院) Peking University(北京大学)

专题命中 安全评测 :safety(abstract)

Comments 10 pages, 4 figures, accepted by ICCVW Author list updated to match the camera-ready version, in compliance with conference policy

详情

展开后加载摘要…

URL PDF HTML 收藏

2. AI治理与伦理 4 篇

2508.12754 2025-08-19 cs.AI 79%

Beyond Ethical Alignment: Evaluating LLMs as Artificial Moral Assistants

Alessio Galatolo, Luca Alberto Rappuoli, Katie Winkle, Meriem Beloucif

机构 * Uppsala University(乌普萨拉大学) University of St. Andrews(圣安德鲁大学)

专题命中 AI治理与伦理 :alignment(title,abstract);分类 cs.AI

Comments Full version of the paper published in ECAI 2025 proceedings (IOS Press, CC BY-NC 4.0)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12174 2025-08-19 cs.CY 57%

Urban AI Governance Must Embed Legal Reasonableness for Democratic and Sustainable Cities

Rashid Mushkani

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11824 2025-08-19 cs.SE cs.AI cs.CR cs.PF 57%

Rethinking Autonomy: Preventing Failures in AI-Driven Software Engineering

Satyam Kumar Navneet, Joydeep Chandra

机构 * Department of CSE Chandigarh University Mohali, India(计算机科学与工程系 奇纳格里大学 莫哈利,印度) Department of CST Tsinghua University Beijing, China(计算机科学与技术系 清华大学 北京,中国)

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10286 2025-08-19 cs.HC 50%

Artificial Emotion: A Survey of Theories and Debates on Realising Emotion in Artificial Intelligence

Yupei Li, Qiyang Sun, Michelle Schlicher, Yee Wen Lim, Björn W. Schuller

专题命中 AI治理与伦理 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 其他安全 14 篇

2508.12803 2025-08-19 cs.CL 79%

When Alignment Hurts: Decoupling Representational Spaces in Multilingual Models

Ahmed Elshabrawy, Hour Kaing, Haiyue Song, Alham Fikri Aji, Hideki Tanaka, Masao Utiyama, Raj Dabre

机构 * MBZUAI(马克斯·普朗克人工智能研究所) NICT, Japan(日本信息通信技术研究所) IIT Madras(印度理工学院Madras分校)

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12422 2025-08-19 cs.CV 67%

Illusions in Humans and AI: How Visual Perception Aligns and Diverges

Jianyi Yang, Junyi Ye, Ankan Dash, Guiling Wang

专题命中 其他安全 :alignment(abstract);safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.01911 2025-08-19 cs.AI cs.CL cs.HC physics.comp-ph 62%

Advancing AI-Scientist Understanding: Multi-Agent LLMs with Interpretable Physics Reasoning

Yinggan Xu, Hana Kimlee, Yijia Xiao, Di Luo

机构 * NSF Center for Quantum Network(NSF量子网络中心) University of California, Los Angeles(加州大学洛杉矶分校)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

Comments ICML 2025 Workshop on MAS

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.06141 2025-08-19 cs.LG cs.AI 62%

Emergent Symbol-like Number Variables in Artificial Neural Networks

Satchel Grant, Noah D. Goodman, James L. McClelland

机构 * Stanford University(斯坦福大学)

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

Journal ref Transactions on Machine Learning Research (TMLR) 2835-8856 (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12872 2025-08-19 cs.DB cs.CY 57%

Evaluating the Quality of Open Building Datasets for Mapping Urban Inequality: A Comparative Analysis Across 5 Cities

Franz Okyere, Meng Lu, Ansgar Brunn

专题命中 其他安全 :alignment(abstract);分类 cs.CY

Comments 25 pages, 4 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10935 2025-08-19 cs.CV cs.LG cs.RO 57%

HQ-OV3D: A High Box Quality Open-World 3D Detection Framework based on Diffision Model

Qi Liu, Yabei Li, Hongsong Wang, Lei He

专题命中 其他安全 :alignment(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.12488 2025-08-19 cs.HC cs.AI 57%

Co-Writing with AI, on Human Terms: Aligning Research with User Demands Across the Writing Process

Mohi Reza, Jeb Thomas-Mitchell, Peter Dushniku, Nathan Laundry, Joseph Jay Williams, Anastasia Kuzminykh

机构 * University of Toronto(多伦多大学)

专题命中 其他安全 :alignment(abstract);分类 cs.AI

Journal ref PACMHCI (CSCW 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.11280 2025-08-19 cs.CL 57%

High-Dimensional Interlingual Representations of Large Language Models

Bryan Wilie, Samuel Cahyawijaya, Junxian He, Pascale Fung

机构 * Hong Kong University of Science and Technology(香港理工大学) Cohere

专题命中 其他安全 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.03885 2025-08-19 cs.LG 57%

Seldonian Reinforcement Learning for Ad Hoc Teamwork

Edoardo Zorzi, Alberto Castellini, Leonidas Bakopoulos, Georgios Chalkiadakis, Alessandro Farinelli

专题命中 其他安全 :safety(abstract);分类 cs.LG

Comments Presented at the 2nd Reinforcement Learning Conference (RLC2025), Edmonton, Canada. To be published in the Proceedings of the Reinforcement Learning Journal 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11889 2025-08-19 cs.CL 57%

In-Context Examples Matter: Improving Emotion Recognition in Conversation with Instruction Tuning

Hui Ma, Bo Zhang, Jinpeng Hu, Zenglin Shi

专题命中 其他安全 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12347 2025-08-19 cs.AR 50%

An ECC-based Fault Tolerance Approach for DNNs

Mohsen Raji, Mohammad Zaree, Kimia Soroush

专题命中 其他安全 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12137 2025-08-19 cs.CV 50%

Infusing fine-grained visual knowledge to Vision-Language Models

Nikolaos-Antonios Ypsilantis, Kaifeng Chen, André Araujo, Ondřej Chum

机构 * VRG, FEE, Czech Technical University in Prague(捷克布拉格技术大学)

专题命中 其他安全 :alignment(abstract)

Comments ICCVW 2025 accepted paper. Workshop name: "What is Next in Multimodal Foundation Models?"

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.09693 2025-08-19 cs.CV 50%

ExpStar: Towards Automatic Commentary Generation for Multi-discipline Scientific Experiments

Jiali Chen, Yujie Jia, Zihan Wu, Jinyu Yang, Jianpeng Chen, Xusen Hei, Jiayuan Xie, Yi Cai, Qing Li

机构 * South China University of Technology(南方科技大学) The Hong Kong Polytechnic University(香港理工大学) Key Laboratory of Big Data and Intelligent Robot Ministry of Education(教育部大数据与智能机器人重点实验室)

专题命中 其他安全 :safety(abstract)

Comments Accepted by ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.04354 2025-08-19 cs.IT eess.SP math.AG math.IT 50%

A transversality theorem for semi-algebraic sets with application to signal recovery from the second moment and cryo-EM

Tamir Bendory, Nadav Dym, Dan Edidin, Arun Suresh

专题命中 其他安全 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏