arXivDaily arXiv每日学术速递 周一至周五更新
爬虫太凶残了,网站这两天有点不稳定,紧急升级ing

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-07-22 至 2025-07-22 共收录 57 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 18 篇

2504.19027 2025-07-22 cs.AI cs.LG cs.NE 62%

DiCE-Extended: A Robust Approach to Counterfactual Explanations in Machine Learning

Volkan Bakir, Polat Goktas, Sureyya Akyuz

机构 * Faculty of Graduate Education Institute, Department of Artificial Intelligence (Interdisciplinary), Bahçeşehir University, Turkey(研究生教育学院人工智能系(跨学科)巴塞希尔大学,土耳其) School of Computer Science, University College Dublin, Ireland(计算机科学学院都柏林大学学院,爱尔兰) Faculty of Engineering and Natural Sciences, Department of Mathematics, Bahçeşehir University, Turkey(工程与自然科学学院数学系巴塞希尔大学,土耳其)

专题命中 安全评测 :alignment(abstract);分类 cs.AI、cs.LG

Comments 5th international Conference on Modelling, Computation and Optimization in Information Systems and Management Sciences (MCO 2025), June 4-6, 2025, Metz, France

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.15761 2025-07-22 cs.AI 57%

GasAgent: A Multi-Agent Framework for Automated Gas Optimization in Smart Contracts

Jingyi Zheng, Zifan Peng, Yule Liu, Junfeng Wang, Yifan Liao, Wenhan Dong, Xinlei He

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.15239 2025-07-22 cs.AI eess.SP 57%

Explainable Artificial Intelligence based Soft Evaluation Indicator for Arc Fault Diagnosis

Qianchao Wang, Yuxuan Ding, Chuanzhen Jia, Zhe Li, Yaping Du

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.15198 2025-07-22 cs.CL 57%

Collaborative Distillation Strategies for Parameter-Efficient Language Model Deployment

Xiandong Meng, Yan Wu, Yexin Tian, Xin Hu, Tianze Kang, Junliang Du

专题命中 安全评测 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.07057 2025-07-22 cs.CL 57%

Tokenization Standards for Linguistic Integrity: Turkish as a Benchmark

M. Ali Bayram, Ali Arda Fincan, Ahmet Semih Gümüş, Sercan Karakaş, Banu Diri, Savaş Yıldırım

机构 * Yıldız Technical University(伊兹密尔技术大学) Yeditepe University(耶迪特佩大学) University of Chicago(芝加哥大学) Istanbul Bilgi University(伊斯坦布尔比金大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14807 2025-07-22 cs.CV cs.AI 57%

Seeing Through Deepfakes: A Human-Inspired Framework for Multi-Face Detection

Juan Hu, Shaojing Fan, Terence Sim

机构 * National University of Singapore(新加坡国立大学)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14513 2025-07-22 cs.AI 57%

Amico: An Event-Driven Modular Framework for Persistent and Embedded Autonomy

Hongyi Yang, Yue Pan, Jiayi Xu, Kelsen Liu

机构 * Department of Aeronautics and Astronautics, Zhejiang University(浙江大学航空航天学院) Department of Computer Science, University College London(伦敦大学学院计算机系) Steinhardt School of Culture, Education, and Human Development, New York University(纽约大学文化、教育与人类发展学院)

专题命中 安全评测 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14492 2025-07-22 cs.LG stat.ML 57%

Glitches in Decision Tree Ensemble Models

Satyankar Chandra, Ashutosh Gupta, Kaushik Mallik, Krishna Shankaranarayanan, Namrita Varshney

专题命中 安全评测 :trustworthy(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14355 2025-07-22 cs.CL 57%

Can LLMs Infer Personality from Real World Conversations?

Jianfeng Zhu, Ruoming Jin, Karin G. Coifman

专题命中 安全评测 :alignment(abstract);分类 cs.CL

Comments 21 pages, 12 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.15321 2025-07-22 cs.CV 50%

BenchDepth: Are We on the Right Way to Evaluate Depth Foundation Models?

Zhenyu Li, Haotong Lin, Jiashi Feng, Peter Wonka, Bingyi Kang

机构 * KAUST(王国塔大学) ByteDance Seed(字节跳动种子) Zhejiang University(浙江大学)

专题命中 安全评测 :alignment(abstract)

Comments Webpage: https://zhyever.github.io/benchdepth

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.11798 2025-07-22 cs.CR 50%

BackdoorDM: A Comprehensive Benchmark for Backdoor Learning on Diffusion Model

Weilin Lin, Nanjun Zhou, Yanyun Wang, Jianze Li, Hui Xiong, Li Liu

专题命中 安全评测 :trustworthy(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14426 2025-07-22 cs.CV 50%

CRAFT: A Neuro-Symbolic Framework for Visual Functional Affordance Grounding

Zhou Chen, Joe Lin, Sathyanarayanan N. Aakur

机构 * Auburn University(亚伯拉罕大学)

专题命中 安全评测 :trustworthy(abstract)

Comments Accepted to NeSy 2025

详情

展开后加载摘要…

URL PDF HTML 收藏

2. AI治理与伦理 2 篇

2507.14339 2025-07-22 cs.CY cs.AI cs.HC cs.LG eess.SP 67%

Fiduciary AI for the Future of Brain-Technology Interactions

Abhishek Bhattacharjee, Jack Pilkington, Nita Farahany

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY、cs.LG

Comments 32 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14332 2025-07-22 cs.LG 57%

Development and Deployment of Hybrid ML Models for Critical Heat Flux Prediction in Annulus Geometries

Aidan Furlong, Xingang Zhao, Robert Salko, Xu Wu

机构 * Department of Nuclear Engineering, North Carolina State University(核工程系,北卡罗来纳州立大学) Department of Nuclear Engineering, University of Tennessee(核工程系,田纳西大学) Nuclear Energy and Fuel Cycle Division, Oak Ridge National Laboratory(核能与燃料循环 division,橡树岭国家实验室)

专题命中 AI治理与伦理 :safety(abstract);分类 cs.LG

Comments Accepted for inclusion in Transactions of the American Nuclear Society for the 2025 ANS Winter Conference

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 其他安全 13 篇

2502.15639 2025-07-22 cs.CL cs.AI cs.LG 82%

Steering into New Embedding Spaces: Analyzing Cross-Lingual Alignment Induced by Model Interventions in Multilingual Language Models

Anirudh Sundar, Sinead Williamson, Katherine Metcalf, Barry-John Theobald, Skyler Seto, Masha Fedzechkina

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.LG

Comments 34 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.15743 2025-07-22 cs.AI cs.CL cs.HC cs.LG 67%

Towards physician-centered oversight of conversational diagnostic AI

Elahe Vedadi, David Barrett, Natalie Harris, Ellery Wulczyn, Shashir Reddy, Roma Ruparel, Mike Schaekermann, Tim Strother, Ryutaro Tanno, Yash Sharma, Jihyeon Lee, Cían Hughes, Dylan Slack, Anil Palepu, Jan Freyberg, Khaled Saab, Valentin Liévin, Wei-Hung Weng, Tao Tu, Yun Liu, Nenad Tomasev, Kavita Kulkarni, S. Sara Mahdavi, Kelvin Guu, Joëlle Barral, Dale R. Webster, James Manyika, Avinatan Hassidim, Katherine Chou, Yossi Matias, Pushmeet Kohli, Adam Rodman, Vivek Natarajan, Alan Karthikesalingam, David Stutz

机构 * Google DeepMind(谷歌DeepMind) Google Research(谷歌研究) Harvard Medical School, Beth Israel Deaconess Medical Center(哈佛医学院,贝塞斯达德acons医学中心)

专题命中 其他安全 :safety(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.08622 2025-07-22 cs.AI cs.CL cs.CV 62%

Visually Guided Decoding: Gradient-Free Hard Prompt Inversion with Language Models

Donghoon Kim, Minji Bae, Kyuhong Shim, Byonghyo Shim

机构 * Seoul National University(首尔国立大学) Sungkyunkwan University(庆熙大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

Comments ICLR 2025 (Official Code: https://github.com/DonghoonKim-1938/VGD)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.15851 2025-07-22 cs.AI 57%

The Other Mind: How Language Models Exhibit Human Temporal Cognition

Lingyu Li, Yang Yao, Yixu Wang, Chubo Li, Yan Teng, Yingchun Wang

专题命中 其他安全 :alignment(abstract);分类 cs.AI

Comments 12 pages, 9 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.10406 2025-07-22 cs.CV cs.AI 57%

RealGeneral: Unifying Visual Generation via Temporal In-Context Learning with Video Models

Yijing Lin, Mengqi Huang, Shuhan Zhuang, Zhendong Mao

机构 * University of Science and Technology of China(中国科学技术大学)

专题命中 其他安全 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.04510 2025-07-22 cs.SE cs.AI 57%

CGP-Tuning: Structure-Aware Soft Prompt Tuning for Code Vulnerability Detection

Ruijun Feng, Hammond Pearce, Pietro Liguori, Yulei Sui

机构 * School of Computer Science and Engineering, University of New South Wales (UNSW)(新南威尔士大学计算机科学与工程学院) Department of Electrical Engineering and Information Technology, University of Naples Federico II(那不勒斯费德里克二世大学电气工程与信息技术系)

专题命中 其他安全 :alignment(abstract);分类 cs.AI

Comments Accepted by IEEE Transactions on Software Engineering

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.01639 2025-07-22 cs.LG cs.SE 57%

ModelVerification.jl: a Comprehensive Toolbox for Formally Verifying Deep Neural Networks

Tianhao Wei, Hanjiang Hu, Luca Marzari, Kai S. Yun, Peizhi Niu, Xusheng Luo, Changliu Liu

机构 * Carnegie Mellon University(卡内基梅隆大学) University of Verona(威尼斯大学)

专题命中 其他安全 :safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.14260 2025-07-22 physics.optics cs.LG 57%

Automating Experimental Optics with Sample Efficient Machine Learning Methods

Arindam Saha, Baramee Charoensombutamon, Thibault Michel, V. Vijendran, Lachlan Walker, Akira Furusawa, Syed M. Assad, Ben C. Buchler, Ping Koy Lam, Aaron D. Tranter

机构 * Australian National University(澳大利亚国立大学) University of Tokyo(东京大学) Agency for Science, Technology and Research(科技研究局) pi Software(2pi软件)

专题命中 其他安全 :alignment(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.16824 2025-07-22 cs.CV cs.AI 57%

PerspectiveNet: Multi-View Perception for Dynamic Scene Understanding

Vinh Nguyen

机构 * Uppsala University(乌普萨拉大学) Florida Institute of Technology(佛罗里达理工学院)

专题命中 其他安全 :safety(abstract);分类 cs.AI

Comments 6 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.15147 2025-07-22 cs.LO cs.FL cs.MA 50%

STL-GO: Spatio-Temporal Logic with Graph Operators for Distributed Systems with Multiple Network Topologies

Yiqi Zhao, Xinyi Yu, Bardh Hoxha, Georgios Fainekos, Jyotirmoy V. Deshmukh, Lars Lindemann

专题命中 其他安全 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14146 2025-07-22 eess.SP 50%

Estimating Markers of Driving Stress through Multimodal Physiological Monitoring

Kleanthis Avramidis, Emily Zhou, Tiantian Feng, Hossein Hamidi Shishavan, Frederico Marcolino Quintao Severgnini, Danny J. Lohan, Paul Schmalenberg, Ercan M. Dede, Shrikanth Narayanan

专题命中 其他安全 :safety(abstract)

Comments 11 pages, 7 figures, 3 tables. This work has been submitted to the IEEE for possible publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.09442 2025-07-22 cs.CV 50%

Advancing Textual Prompt Learning with Anchored Attributes

Zheng Li, Yibing Song, Ming-Ming Cheng, Xiang Li, Jian Yang

机构 * PCA Lab, VCIP, College of Computer Science, Nankai University(PCA实验室、VCIP、计算机科学学院、南开大学) DAMO Academy, Alibaba Group(达摩院、阿里巴巴集团)

专题命中 其他安全 :alignment(abstract)

Comments ICCV 2025. Code: https://github.com/zhengli97/ATPrompt. Project Page: https://zhengli97.github.io/ATPrompt/

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.18507 2025-07-22 cs.RO 50%

At First Contact: Stiffness Estimation Using Vibrational Information for Prosthetic Grasp Modulation

Anway S. Pimpalkar, Ariel Slepyan, Nitish V. Thakor

机构 * Department of Biomedical Engineering, Johns Hopkins University(生物医学工程系,约翰·霍普金斯大学) Department of Electrical and Computer Engineering, Johns Hopkins University(电气与计算机工程系,约翰·霍普金斯大学)

专题命中 其他安全 :safety(abstract)

Comments 5 pages, 7 figures, for IEEE Sensors Letters

详情

展开后加载摘要…

URL PDF HTML 收藏