arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-08-26 至 2025-08-26 共收录 77 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 幻觉与事实性 6 篇

2508.16974 2025-08-26 cs.CV 50%

Hierarchical Contextual Grounding LVLM: Enhancing Fine-Grained Visual-Language Understanding with Robust Grounding

Leilei Guo, Antonio Carlos Rivera, Peiyu Tang, Haoxuan Ren, Zheyu Song

机构 * Zhongkai University of Agriculture and Engineering(仲恺农业工程大学) EDP University of Puerto Rico: San Sebastian(波多黎各圣塞巴斯蒂安EDP大学)

专题命中 幻觉与事实性 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 隐私与版权 3 篇

2411.16796 2025-08-26 cs.LG cs.CL cs.CV cs.DC 62%

HeteroTune: Efficient Federated Learning for Large Heterogeneous Models

Ruofan Jia, Weiying Xie, Jie Lei, Jitao Ma, Haonan Qin, Leyuan Fang

专题命中 隐私与版权 :alignment(abstract);分类 cs.CL、cs.LG

Comments 16 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.16713 2025-08-26 cs.SE cs.AI hep-ex 57%

CelloAI: Leveraging Large Language Models for HPC Software Development in High Energy Physics

Mohammad Atif, Kriti Chopra, Ozgur Kilic, Tianle Wang, Zhihua Dong, Charles Leggett, Meifeng Lin, Paolo Calafiura, Salman Habib

机构 * Brookhaven National Laboratory(布鲁克海文国家实验室) Lawrence Berkeley National Laboratory(伯克利国家实验室) Argonne National Laboratory(阿贡国家实验室)

专题命中 隐私与版权 :safety(abstract);分类 cs.AI

Comments 12 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17653 2025-08-26 cs.CV 50%

FloraSyntropy-Net: Scalable Deep Learning with Novel FloraSyntropy Archive for Large-Scale Plant Disease Diagnosis

Saif Ur Rehman Khan, Muhammad Nabeel Asim, Sebastian Vollmer, Andreas Dengel

机构 * German Research Center for Artificial Intelligence(德国人工智能研究中心) Intelligentx GmbH

专题命中 隐私与版权 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 安全评测 24 篇

2508.11398 2025-08-26 cs.HC cs.AI cs.IR 79%

Trustworthy AI Psychotherapy: Multi-Agent LLM Workflow for Counseling and Explainable Mental Disorder Diagnosis

Mithat Can Ozgun, Jiahuan Pei, Koen Hindriks, Lucia Donatelli, Qingzhi Liu, Junxiao Wang

机构 * Vrije Universiteit Amsterdam(阿姆斯特丹自由大学) Wageningen University and Research(瓦赫宁根大学和研究中心) Guangzhou University(广州大学)

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI

Comments This paper has been accepted by CIKM 2025 as a full paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.14527 2025-08-26 cs.CV 78%

Adversarial Generation and Collaborative Evolution of Safety-Critical Scenarios for Autonomous Vehicles

Jiangfan Liu, Yongkang Guo, Fangzhi Zhong, Tianyuan Zhang, Zonglei Jing, Siyuan Liang, Jiakai Wang, Mingchuan Zhang, Aishan Liu, Xianglong Liu

机构 * Beihang University(北京航空航天大学) Nanyang Technological University(南洋理工大学) Zhongguancun Laboratory(中关村实验室) Henan University of Science and Technology(河南科技大学)

专题命中 安全评测 :safety(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.16850 2025-08-26 cs.AI 70%

RADAR: A Reasoning-Guided Attribution Framework for Explainable Visual Data Analysis

Anku Rani, Aparna Garimella, Apoorv Saxena, Balaji Vasan Srinivasan, Paul Pu Liang

专题命中 安全评测 :alignment(abstract);trustworthy(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.14377 2025-08-26 cs.CL cs.AI cs.CY 67%

ZPD-SCA: Unveiling the Blind Spots of LLMs in Assessing Students' Cognitive Abilities

Wenhan Dong, Zhen Sun, Yuemeng Zhao, Zifan Peng, Jun Wu, Jingyi Zheng, Yule Liu, Xinlei He, Yu Wang, Ruiming Wang, Xinyi Huang, Lei Mo

机构 * School of Psychology, South China Normal University(南方科技大学心理学院) Information Hub, Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)信息中心) School of AI, Guangzhou University(广州大学人工智能学院) College of Cyber Security, Jinan University(济南大学网络安全学院)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17786 2025-08-26 cs.AI cs.FL cs.LG cs.LO 62%

Interpretable Early Failure Detection via Machine Learning and Trace Checking-based Monitoring

Andrea Brunello, Luca Geatti, Angelo Montanari, Nicola Saccomanno

机构 * University of Udine, Italy(乌迪内大学)

专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.LG

Comments Full version of the paper accepted for publication at the 28th European Conference on Artificial Intelligence (ECAI 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17680 2025-08-26 cs.LG cs.AI cs.CV 62%

Robustness Feature Adapter for Efficient Adversarial Training

Quanwei Wu, Jun Guo, Wei Wang, Yi Wang

机构 * Dongguan University of Technology(东莞科技学院) The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州))

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG

Comments The paper has been accepted for presentation at ECAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11280 2025-08-26 cs.CL cs.AI 62%

LETToT: Label-Free Evaluation of Large Language Models On Tourism Using Expert Tree-of-Thought

Ruiyan Qi, Congding Wen, Weibo Zhou, Jiwei Li, Shangsong Liang, Lingbo Li

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.14745 2025-08-26 cs.LG cs.AI 62%

Explainable Prediction of the Mechanical Properties of Composites with CNNs

Varun Raaghav, Dimitrios Bikos, Antonio Rago, Francesca Toni, Maria Charalambides

机构 * Department of Mechanical Engineering, Imperial College London, UK(帝国理工学院机械工程系) Department of Computing, Imperial College London, UK(帝国理工学院计算系) Department of Informatics, King's College London, UK(伦敦国王学院信息学系)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG

Comments 9 pages, 6 figures. Accepted for publication at The 14th Conference on Prestigious Applications of Intelligent Systems (PAIS-2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.16672 2025-08-26 cs.CY cs.AI 62%

The AI Model Risk Catalog: What Developers and Researchers Miss About Real-World AI Harms

Pooja S. B. Rao, Sanja Šćepanović, Dinesh Babu Jayagopi, Mauro Cherubini, Daniele Quercia

专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.CY

Comments Accepted to AIES 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.16669 2025-08-26 cs.CY cs.AI cs.HC 62%

Situational Awareness as the Imperative Capability for Disaster Resilience in the Era of Complex Hazards and Artificial Intelligence

Hongrak Pak, Ali Mostafavi

机构 * organization= UrbanResilience.AI Lab, Zachry Department of Civil Environmental Engineering, Texas A\&M University , city= College Station , postcode= 77843 , state= TX , country= USA

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17667 2025-08-26 cs.CV cs.AI 57%

Hierarchical Vision-Language Learning for Medical Out-of-Distribution Detection

Runhe Lai, Xinhua Lu, Kanghao Chen, Qichao Chen, Wei-Shi Zheng, Ruixuan Wang

机构 * Peng Cheng Laboratory(鹏城实验室) Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) University of Nottingham Malaysia(诺丁汉大学(马来西亚)) Key Laboratory of Machine Intelligence and Advanced Computing, MOE(机器智能与高级计算重点实验室)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

Comments 10 pages, 2 figures, Accepted by MICCAI2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17405 2025-08-26 cs.LG cs.CR 57%

FRAME : Comprehensive Risk Assessment Framework for Adversarial Machine Learning Threats

Avishag Shapira, Simon Shigol, Asaf Shabtai

机构 * Ben-Gurion University of the Negev(贝内杰尔大学)

专题命中 安全评测 :alignment(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17378 2025-08-26 cs.CL 57%

UI-Level Evaluation of ALLaM 34B: Measuring an Arabic-Centric LLM via HUMAIN Chat

Omer Nacar

机构 * NAMAA Community(NAMAA社区) Riyadh - KSA(利雅得-科威特)

专题命中 安全评测 :safety(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.09329 2025-08-26 cs.AI cs.CR 57%

When Developer Aid Becomes Security Debt: A Systematic Analysis of Insecure Behaviors in LLM Coding Agents

Matous Kozak, Roshanak Zilouchian Moghaddam, Siva Sivaraman

机构 * Microsoft(微软公司) Czech Technical University in Prague(捷克技术大学)

专题命中 安全评测 :safety(abstract);分类 cs.AI

Comments 15 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.00258 2025-08-26 cs.AI 57%

Hidden in Plain Sight: Reasoning in Underspecified and Misspecified Scenarios for Multimodal LLMs

Qianqi Yan, Hongquan Li, Shan Jiang, Yang Zhao, Xinze Guan, Ching-Chen Kuo, Xin Eric Wang

机构 * University of California, Santa Cruz(加州大学圣克鲁兹分校) eBay

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.06791 2025-08-26 cs.RO cs.AI cs.HC cs.MA 57%

AutoMisty: A Multi-Agent LLM Framework for Automated Code Generation in the Misty Social Robot

Xiao Wang, Lu Dong, Sahana Rangasrinivasan, Ifeoma Nwogu, Srirangaraj Setlur, Venugopal Govindaraju

机构 * State University of New York at Buffalo(纽约州立大学布法罗分校) Department of Computer Science and Engineering, Amrita School of Computing, Amrita Vishwa Vidyapeetham(计算机科学与工程系,阿米特拉学校 computing,阿米特拉世界大学)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

Comments Accepted by IROS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.01027 2025-08-26 stat.ML cs.LG 57%

Adversarial Robustness in Two-Stage Learning-to-Defer: Algorithms and Guarantees

Yannis Montreuil, Axel Carlier, Lai Xing Ng, Wei Tsang Ooi

机构 * School of Computing, National University of Singapore, Singapore(新加坡国立大学计算机学院) IRIT, Université de Toulouse, CNRS, Toulouse INP, Toulouse, France(图卢兹大学IRIT研究所、法国国家科学研究中心、图卢兹INP) Institute for Infocomm Research, Agency for Science, Technology and Research, Singapore(新加坡科技研究局信息与通信研究所)

专题命中 安全评测 :trustworthy(abstract);分类 cs.LG

Comments Accepted at the 42nd International Conference on Machine Learning (ICML 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2202.13852 2025-08-26 cs.LG 57%

Hyperbolic Graph Neural Networks: A Review of Methods and Applications

Menglin Yang, Min Zhou, Tong Zhang, Jiahong Liu, Zhihao Li, Lujia Pan, Hui Xiong, Irwin King

机构 * Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) Huawei Technologies Co., Ltd.(华为技术有限公司) The Chinese University of Hong Kong(香港中文大学) Zhejiang University(浙江大学)

专题命中 安全评测 :trustworthy(abstract);分类 cs.LG

Comments The latest draft was circulated under the title "Hyperbolic Graph Learning: A Comprehensive Review." The present arXiv version retains the original title, "Hyperbolic Graph Neural Networks: A Review of Methods and Applications," for consistency, while incorporating substantial revisions and extensions

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.15075 2025-08-26 cs.CL cs.AI cs.CV cs.LG 56%

Traveling Across Languages: Benchmarking Cross-Lingual Consistency in Multimodal LLMs

Hao Wang, Pinzhi Huang, Jihan Yang, Saining Xie, Daisuke Kawahara

专题命中 安全评测 :分类 cs.CL、cs.AI、cs.LG;prompt injection(comments)

Comments The first version of this paper mistakenly included a prompt injection phrase, which was inappropriate and unprofessional. Although we corrected the version on arXiv and withdrew from the conference, my co-authors and university strongly request a full withdrawal. Given the situation, I no longer have the authority to manage this paper, and withdrawing it from arXiv is the most responsible action

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17975 2025-08-26 cs.CV math.LO 50%

Enhanced Drift-Aware Computer Vision Architecture for Autonomous Driving

Md Shahi Amran Hossain, Abu Shad Ahammed, Sayeri Mukherjee, Roman Obermaisser

机构 * Chair of Embedded Systems University of Siegen(嵌入系统教授会 University of Siegen)

专题命中 安全评测 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17255 2025-08-26 cs.CV cs.RO 50%

SEER-VAR: Semantic Egocentric Environment Reasoner for Vehicle Augmented Reality

Yuzhi Lai, Shenghai Yuan, Peizheng Li, Jun Lou, Andreas Zell

专题命中 安全评测 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.13491 2025-08-26 cs.CE 50%

From Scores to Skills: A Cognitive Diagnosis Framework for Evaluating Financial Large Language Models

Ziyan Kuang, Feiyu Zhu, Maowei Jiang, Yanzhao Lai, Zelin Wang, Zhitong Wang, Meikang Qiu, Jiajia Huang, Min Peng, Qianqian Xie, Sophia Ananiadou

专题命中 安全评测 :trustworthy(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.16661 2025-08-26 cs.CV 50%

QA-VLM: Providing human-interpretable quality assessment for wire-feed laser additive manufacturing parts with Vision Language Models

Qiaojie Zheng, Jiucai Zhang, Joy Gockel, Michael B. Wakin, Craig Brice, Xiaoli Zhang

机构 * Mechanical Engineering, Colorado School of Mines(机械工程,科罗拉多矿业学院) Electrical Engineering, Colorado School of Mines(电气工程,科罗拉多矿业学院)

专题命中 安全评测 :trustworthy(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.18694 2025-08-26 cs.SE 50%

Requirements-Driven Automated Software Testing: A Systematic Review

Fanyu Wang, Chetan Arora, Chakkrit Tantithamthavorn, Kaicheng Huang, Aldeida Aleti

专题命中 安全评测 :alignment(abstract)

Comments Accepted by TOSEM

详情

展开后加载摘要…

URL PDF HTML 收藏

4. AI治理与伦理 2 篇

2508.16762 2025-08-26 cs.CL cs.CY 62%

Toward Socially Aware Vision-Language Models: Evaluating Cultural Competence Through Multimodal Story Generation

Arka Mukherjee, Shreya Ghosh

机构 * KIIT Deemed University(KIIT大学) Indian Institute of Technology (IIT) Bhubaneswar(印度理工学院(Bhubaneswar分校))

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.CY

Comments Accepted at ASI @ ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.16642 2025-08-26 cs.CY 57%

AI as IA: The use and abuse of artificial intelligence (AI) for human enhancement through intellectual augmentation (IA)

Alexandre Erler, Vincent C. Müller

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CY

Journal ref (2023) in Marcello Ienca and Fabrice Jotterand (eds.), The Routledge Handbook of the Ethics of Human Enhancement (London: Routledge), 187-99

详情

展开后加载摘要…

URL PDF HTML 收藏