arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 9451 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 9451 篇

2507.10808 2025-10-07 cs.CR cs.SY eess.SP eess.SY 50%

Contrastive-KAN: A Semi-Supervised Intrusion Detection Framework for Cybersecurity with scarce Labeled Data

Mohammad Alikhani, Reza Kazemi

专题命中 安全评测 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02994 2025-10-06 cs.CV 50%

Towards Scalable and Consistent 3D Editing

Ruihao Xia, Yang Tang, Pan Zhou

机构 * East China University of Science and Technology(东华大学) Singapore Management University(新加坡管理学院)

专题命中 安全评测 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02909 2025-10-06 cs.CV 50%

Training-Free Out-Of-Distribution Segmentation With Foundation Models

Laith Nayal, Hadi Salloum, Ahmad Taha, Yaroslav Kholodov, Alexander Gasnikov

机构 * Laboratory of Multimodal Research In Industry, AI Institute, Innopolis University(工业多模态研究实验室,人工智能研究所,因诺普利斯大学) Phystech School of Applied Mathematics and Computer Science, Moscow Institute of Physics and Technology(物理与技术莫斯科应用数学与计算机科学学院,莫斯科物理技术学院) Research Center for Artificial Intelligence, Innopolis University(人工智能研究中心,因诺普利斯大学) Q Deep, Innopolis(Q深度,因诺普利斯) Machine Learning and Data Representation Lab, Innopolis University(机器学习与数据表示实验室,因诺普利斯大学)

专题命中 安全评测 :safety(abstract)

Comments 12 pages, 5 figures, 2 tables, ICOMP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02722 2025-10-06 cs.CV 50%

MoGIC: Boosting Motion Generation via Intention Understanding and Visual Context

Junyu Shi, Yong Sun, Zhiyuan Zhang, Lijiang Liu, Zhengjie Zhang, Yuxin He, Qiang Nie

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))

专题命中 安全评测 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20088 2025-10-06 cs.CV cs.MM cs.SD 50%

AudioStory: Generating Long-Form Narrative Audio with Large Language Models

Yuxin Guo, Teng Wang, Yuying Ge, Shijie Ma, Yixiao Ge, Wei Zou, Ying Shan

机构 * School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) ARC Lab, Tencent PCG(腾讯PCG ARC实验室) MAIS, Institute of Automation, CAS, Beijing(自动化研究所北京研究所MAIS)

专题命中 安全评测 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.01534 2025-10-06 cs.CV 50%

Toward a Holistic Evaluation of Robustness in CLIP Models

Weijie Tu, Weijian Deng, Tom Gedeon

专题命中 安全评测 :safety(abstract)

Comments Accepted to IEEE TPAMI, extension of NeurIPS'23 work: A Closer Look at the Robustness of Contrastive Language-Image Pre-Training (CLIP)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02169 2025-10-03 cs.SE cs.CR 50%

TAIBOM: Bringing Trustworthiness to AI-Enabled Systems

Vadim Safronov, Anthony McCaigue, Nicholas Allott, Andrew Martin

专题命中 安全评测 :trustworthy(abstract)

Comments This paper has been accepted at the First International Workshop on Security and Privacy-Preserving AI/ML (SPAIML 2025), co-located with the 28th European Conference on Artificial Intelligence (ECAI 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01683 2025-10-03 cs.CV 50%

Uncovering Overconfident Failures in CXR Models via Augmentation-Sensitivity Risk Scoring

Han-Jay Shu, Wei-Ning Chiu, Shun-Ting Chang, Meng-Ping Huang, Takeshi Tohyama, Ahram Han, Po-Chih Kuo

机构 * National Tsing Hua University(国立清华大学) National Taiwan University(国立台湾大学) Massachusetts Institute of Technology(麻省理工学院)

专题命中 安全评测 :safety(abstract)

Comments 5 pages, 1 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04677 2025-10-03 cs.CV 50%

Robust Prompt Tuning for Vision-Language Models with Mild Semantic Noise

Yansheng Gao, Yufei Zheng, Shengsheng Wang

机构 * College of Computer Science and Technology, Key Laboratory of Symbolic Computation and Knowledge Engineering of Ministry of Education, Jilin University(计算机科学与技术学院、教育部符号计算与知识工程重点实验室、吉林大学) College of Software, Key Laboratory of Symbolic Computation and Knowledge Engineering of Ministry of Education, Jilin University(软件学院、教育部符号计算与知识工程重点实验室、吉林大学)

专题命中 安全评测 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.00932 2025-10-02 cs.PF 50%

Opal: A Modular Framework for Optimizing Performance using Analytics and LLMs

Mohammad Zaeed, Tanzima Z. Islam, Vladimir Inđić

专题命中 安全评测 :trustworthy(abstract)

Comments 12 pages and 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.00583 2025-10-02 cs.HC 50%

Rethinking Wine Tasting for Chinese Consumers: A Service Design Approach Enhanced by Multimodal Personalization

Xinyang Shan, Yuanyuan Xu, Tian Xia, Yinshan Lin

专题命中 安全评测 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.00205 2025-10-02 q-fin.CP 50%

Quantifying Semantic Shift in Financial NLP: Robust Metrics for Market Prediction Stability

Zhongtian Sun, Chenghao Xiao, Anoushka Harit, Jongmin Yu

专题命中 安全评测 :alignment(abstract)

Comments The 6th ACM International Conference on Al in Finance

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.26006 2025-10-02 cs.CV 50%

AgenticIQA: An Agentic Framework for Adaptive and Interpretable Image Quality Assessment

Hanwei Zhu, Yu Tian, Keyan Ding, Baoliang Chen, Bolin Chen, Shiqi Wang, Weisi Lin

机构 * Nanyang Technological University(南洋理工大学) Nanjing University of Information Science and Technology(南京信息工程大学) Zhejiang University(浙江大学) South China Normal University(华南师范大学) Alibaba DAMO Academy(阿里巴巴达摩院) City University of Hong Kong(香港城市大学)

专题命中 安全评测 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.11920 2025-10-02 cs.RO cs.SY eess.SY 50%

Heterogeneous Predictor-based Risk-Aware Planning with Conformal Prediction in Dense, Uncertain Environments

Jeongyong Yang, KwangBin Lee, SooJean Han

机构 * School of Electrical Engineering, Korea Advanced Institute of Science and Technology (KAIST)(电气工程学院,韩国科学技术院)

专题命中 安全评测 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.26223 2025-10-01 q-bio.GN 50%

Nephrobase Cell+: Multimodal Single-Cell Foundation Model for Decoding Kidney Biology

Chenyu Li, Elias Ziyadeh, Yash Sharma, Bernhard Dumoulin, Jonathan Levinsohn, Eunji Ha, Siyu Pan, Vishwanatha Rao, Madhav Subramaniyam, Mario Szegedy, Nancy Zhang, Katalin Susztak

专题命中 安全评测 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25974 2025-10-01 cs.NI cs.MA 50%

OpenID Connect for Agents (OIDC-A) 1.0: A Standard Extension for LLM-Based Agent Identity and Authorization

Subramanya Nagabhushanaradhya

专题命中 安全评测 :trustworthy(abstract)

Comments 10 pages, 5 tables, 2 code listings. Specification proposal available at https://github.com/subramanya1997/oidc-a/

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23991 2025-09-30 cs.CV 50%

RPG360: Robust 360 Depth Estimation with Perspective Foundation Models and Graph Optimization

Dongki Jung, Jaehoon Choi, Yonghan Lee, Dinesh Manocha

机构 * University of Maryland, College Park(马里兰大学 College Park 分校)

专题命中 安全评测 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23251 2025-09-30 cs.MM cs.SD 50%

XGC-AVis: Towards Audio-Visual Content Understanding with a Multi-Agent Collaborative System

Yuqin Cao, Xiongkuo Min, Yixuan Gao, Wei Sun, Zicheng Zhang, Jinliang Han, Guangtao Zhai

机构 * Shanghai Jiao Tong University(上海交通大学) East China Normal University(华东师范大学) Shanghai AI Laboratory(上海人工智能实验室)

专题命中 安全评测 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22228 2025-09-29 cs.CV 50%

UrbanFeel: A Comprehensive Benchmark for Temporal and Perceptual Understanding of City Scenes through Human Perspective

Jun He, Yi Lin, Zilong Huang, Jiacong Yin, Junyan Ye, Yuchuan Zhou, Weijia Li, Xiang Zhang

机构 * Sun Yat-sen University(中山大学)

专题命中 安全评测 :safety(abstract)

Comments 13 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.03334 2025-09-29 cs.CV cs.DB 50%

OS-W2S: An Automatic Labeling Engine for Language-Guided Open-Set Aerial Object Detection

Guoting Wei, Yu Liu, Xia Yuan, Xizhe Xue, Linlin Guo, Yifan Yang, Chunxia Zhao, Zongwen Bai, Haokui Zhang, Rong Xiao

机构 * Nanjing University of Science and Technology(南京理工大学) Intellifusion Inc.(Intellifusion公司) Northwestern Polytechnical University(西北工业大学) Zhejiang Lab(浙江实验室) Yan’an University(延安大学) Beijing University of Posts and Telecommunications(北京邮电大学)

专题命中 安全评测 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19096 2025-09-26 cs.CV cs.SE 50%

Investigating Traffic Accident Detection Using Multimodal Large Language Models

Ilhan Skender, Kailin Tong, Selim Solmaz, Daniel Watzenig

机构 * Embedded Systems Group (Dept.-E)(嵌入式系统组) Virtual Vehicle Research GmbH(虚拟车辆研究公司) Control Systems Group (Dept.-E)(控制系统组) Institute of Visual Computing(视觉计算研究所) Graz University of Technology(格拉茨技术大学)

专题命中 安全评测 :safety(abstract)

Comments Accepted for presentation at the 2025 IEEE International Automated Vehicle Validation Conference (IAVVC 2025). Final version to appear in IEEE Xplore

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.18485 2025-09-25 q-bio.NC cs.CV 50%

Deciphering Functions of Neurons in Vision-Language Models

Jiaqi Xu, Cuiling Lan, Yan Lu

机构 * University of Science and Technology of China(中国科学技术大学) Microsoft Research Asia(微软亚洲研究院)

专题命中 安全评测 :trustworthy(abstract)

Comments Accepted by the 31st ACM International Conference on Multimedia (ACM MM 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19958 2025-09-25 cs.RO 50%

Generalist Robot Manipulation beyond Action Labeled Data

Alexander Spiridonov, Jan-Nico Zaech, Nikolay Nikolov, Luc Van Gool, Danda Pani Paudel

机构 * ETH Zurich, Switzerland(苏黎世联邦理工学院,瑞士)

专题命中 安全评测 :alignment(abstract)

Comments Accepted at Conference on Robot Learning 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19831 2025-09-25 eess.AS 50%

SCORE: Scaling audio generation using Standardized COmposite REwards

Jaemin Jung, Jaehun Kim, Inkyu Shin, Joon Son Chung

专题命中 安全评测 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18869 2025-09-24 cs.DC 50%

On The Reproducibility Limitations of RAG Systems

Baiqiang Wang, Dongfang Zhao, Nathan R Tallent, Luanzheng Guo

专题命中 安全评测 :trustworthy(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17537 2025-09-24 cs.CV 50%

SimToken: A Simple Baseline for Referring Audio-Visual Segmentation

Dian Jin, Yanghao Zhou, Jinxing Zhou, Jiaqi Ma, Ruohao Guo, Dan Guo

专题命中 安全评测 :alignment(abstract)

Comments Project page: https://github.com/DianJin-HFUT/SimToken

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16736 2025-09-23 cs.MA cs.CR 50%

Towards Transparent and Incentive-Compatible Collaboration in Decentralized LLM Multi-Agent Systems: A Blockchain-Driven Approach

Minfeng Qi, Tianqing Zhu, Lefeng Zhang, Ningran Li, Wanlei Zhou

专题命中 安全评测 :trustworthy(abstract)

Comments 17 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15432 2025-09-22 cs.IR 50%

SERVAL: Surprisingly Effective Zero-Shot Visual Document Retrieval Powered by Large Vision and Language Models

Thong Nguyen, Yibin Lei, Jia-Huei Ju, Andrew Yates

专题命中 安全评测 :alignment(abstract)

Comments Accepted

Journal ref EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15068 2025-09-19 cs.HC 50%

Learning in Context: Personalizing Educational Content with Large Language Models to Enhance Student Learning

Joy Jia Yin Lim, Daniel Zhang-Li, Jifan Yu, Xin Cong, Ye He, Zhiyuan Liu, Huiqin Liu, Lei Hou, Juanzi Li, Bin Xu

专题命中 安全评测 :trustworthy(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.14739 2025-09-19 cs.CV 50%

FMGS-Avatar: Mesh-Guided 2D Gaussian Splatting with Foundation Model Priors for 3D Monocular Avatar Reconstruction

Jinlong Fan, Bingyu Hu, Xingguang Li, Yuxiang Yang, Jing Zhang

机构 * HangZhou Dianzi University(杭州电子大学) Shenzhen Polytechnic University(深圳职业技术大学) WuHan University(武汉大学)

专题命中 安全评测 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏