arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-09-10 至 2025-09-10 共收录 13 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 13 篇

2508.17450 2025-09-10 cs.CL cs.CY 84%

Persuasion Dynamics in LLMs: Investigating Robustness and Adaptability in Knowledge and Safety with DuET-PD

Bryan Chen Zhengyu Tan, Daniel Wai Kit Chin, Zhengyuan Liu, Nancy F. Chen, Roy Ka-Wei Lee

机构 * Singapore University of Technology and Design (SUTD)(新加坡科技设计大学) Institute for Infocomm Research (I2R), A*STAR, Singapore(信息通信研究院)

专题命中 安全评测 :safety(title,abstract);DPO(abstract);分类 cs.CL、cs.CY

Comments To appear at EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.07315 2025-09-10 cs.CR cs.SE 82%

SafeToolBench: Pioneering a Prospective Benchmark to Evaluating Tool Utilization Safety in LLMs

Hongfei Xia, Hongru Wang, Zeming Liu, Qian Yu, Yuhang Guo, Haifeng Wang

专题命中 安全评测 :safety(title,abstract);trustworthy(abstract)

Comments 18 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00700 2025-09-10 cs.CV 78%

Prompt the Unseen: Evaluating Visual-Language Alignment Beyond Supervision

Raehyuk Jung, Seungjun Yu, Hyunjung Shim

专题命中 安全评测 :alignment(title,abstract)

Comments Link to publicly available codes is added

详情

展开后加载摘要…

URL PDF HTML 收藏
2101.12100 2025-09-10 cs.LG cs.AI 73%

Increasing the Confidence of Deep Neural Networks by Coverage Analysis

Giulio Rossolini, Alessandro Biondi, Giorgio Buttazzo

专题命中 安全评测 :safety(abstract);trustworthy(abstract);分类 cs.AI、cs.LG

Journal ref IEEE Transactions on Software Engineering ( Volume: 49, Issue: 2, 01 February 2023)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.18933 2025-09-10 cs.AI cs.CR cs.CY cs.LG 67%

VISION: Robust and Interpretable Code Vulnerability Detection Leveraging Counterfactual Augmentation

David Egea, Barproda Halder, Sanghamitra Dutta

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.CY、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.07017 2025-09-10 cs.AI cs.CL cs.LG 67%

From Eigenmodes to Proofs: Integrating Graph Spectral Operators with Symbolic Interpretable Reasoning

Andrew Kiruluta, Priscilla Burity

机构 * Andrew Kiruluta and Priscilla Burity(独立研究者)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.07473 2025-09-10 cs.AI 57%

SheetDesigner: MLLM-Powered Spreadsheet Layout Generation with Rule-Based and Vision-Based Reflection

Qin Chen, Yuanyi Ren, Xiaojun Ma, Mugeng Liu, Han Shi, Dongmei Zhang

机构 * Peking University(北京大学) Microsoft(微软)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

Comments Accepted to EMNLP 2025 Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.07127 2025-09-10 cs.GR cs.AI cs.CV 57%

SVGauge: Towards Human-Aligned Evaluation for SVG Generation

Leonardo Zini, Elia Frigieri, Sebastiano Aloscari, Marcello Generali, Lorenzo Dodi, Robert Dosen, Lorenzo Baraldi

机构 * University of Modena and Reggio Emilia(摩德纳和雷吉奥艾米利亚大学) Doxee S.p.A.(Doxee公司)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

Comments Accepted at 23rd edition of International Conference on Image Analysis and Processing 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05702 2025-09-10 cs.MA cs.AI cs.SY eess.SY 57%

Grid-Agent: An LLM-Powered Multi-Agent System for Power Grid Control

Yan Zhang, Ahmad Mohammad Saber, Amr Youssef, Deepa Kundur

机构 * Department of Electrical and Computer Engineering, University of Toronto(电气与计算机工程系,多伦多大学) Concordia Institute for Information Systems Engineering(康卡迪亚信息系统工程研究所) Concordia University(康卡迪亚大学)

专题命中 安全评测 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.07039 2025-09-10 cs.LG cs.CV 57%

Benchmarking Vision Transformers and CNNs for Thermal Photovoltaic Fault Detection with Explainable AI Validation

Serra Aksoy

专题命中 安全评测 :alignment(abstract);分类 cs.LG

Comments 28 Pages, 4 Figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.07026 2025-09-10 cs.LO cs.AI 57%

Contradictions

Yang Xu, Shuwei Chen, Xiaomei Zhong, Jun Liu, Xingxing He

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

Comments 37 Pages,9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.07010 2025-09-10 cs.CV cs.AI cs.ET 57%

Human-in-the-Loop: Quantitative Evaluation of 3D Models Generation by Large Language Models

Ahmed R. Sadik, Mariusz Bujny

机构 * Honda Research Institute Europe - Germany(本田欧洲研究机构)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.07323 2025-09-10 cs.SD cs.CR 50%

When Fine-Tuning is Not Enough: Lessons from HSAD on Hybrid and Adversarial Audio Spoof Detection

Bin Hu, Kunyang Huang, Daehan Kwak, Meng Xu, Kuan Huang

机构 * Department of Computer Science and Technology, Kean University, USA(计算机科学与技术系,凯恩大学,美国) Department of Computer Science and Technology, Wenzhou-Kean University, China(计算机科学与技术系,温州-凯恩大学,中国)

专题命中 安全评测 :trustworthy(abstract)

Comments 13 pages, 11 figures.This work has been submitted to the IEEE for possible publication

详情

展开后加载摘要…

URL PDF HTML 收藏