arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-09-10 至 2025-09-10 共收录 32 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全训练 3 篇

2503.16833 2025-09-10 cs.SD cs.AI cs.CL cs.CY eess.AS 67%

The Model Hears You: Audio Language Model Deployments Should Consider the Principle of Least Privilege

Luxi He, Xiangyu Qi, Michel Liao, Inyoung Cheong, Prateek Mittal, Danqi Chen, Peter Henderson

专题命中 安全训练 :safety(abstract);分类 cs.CL、cs.AI、cs.CY

Comments Published at AIES 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.07845 2025-09-10 cs.LG 57%

Predicting person-level injury severity using crash narratives: A balanced approach with roadway classification and natural language process techniques

Mohammad Zana Majidi, Sajjad Karimi, Teng Wang, Robert Kluger, Reginald Souleyrette

专题命中 安全训练 :safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.07304 2025-09-10 eess.SY cs.SY 50%

Distributed Leader-Follower Consensus for Uncertain Multiagent Systems with Time-Triggered Switching of the Communication Network

Armel Koulong, Ali Pakniyat

专题命中 安全训练 :safety(abstract)

Comments Joint submission paper MECC-JDSMC. Accepted for the 2025 Modeling, Estimation and Control Conference (MECC). Currently under review by the ASME Journal of Dynamic Systems, Measurement, and Control (JDSMC)

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 越狱攻击 2 篇

2508.17674 2025-09-10 cs.CR cs.AI cs.LG 62%

Attacking LLMs and AI Agents: Advertisement Embedding Attacks Against Large Language Models

Qiming Guo, Jinwen Tang, Xingran Huang

机构 * Department of Computer Science Texas A\&M University–Corpus Christi Corpus Christi, TX, USA EECS Department University of Missouri Columbia, MO, USA Department of Computer Engineering University of California–Riverside Riverside, CA, USA

专题命中 越狱攻击 :safety(abstract);分类 cs.AI、cs.LG

Comments 6 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.07941 2025-09-10 cs.CR cs.AI 57%

ImportSnare: Directed "Code Manual" Hijacking in Retrieval-Augmented Code Generation

Kai Ye, Liangcai Su, Chenxiong Qian

机构 * The University of Hong Kong(香港大学)

专题命中 越狱攻击 :alignment(abstract);分类 cs.AI

Comments This paper has been accepted by the ACM Conference on Computer and Communications Security (CCS) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 提示注入 1 篇

2509.07617 2025-09-10 cs.AI 79%

Transferable Direct Prompt Injection via Activation-Guided MCMC Sampling

Minghui Li, Hao Zhang, Yechao Zhang, Wei Wan, Shengshan Hu, pei Xiaobing, Jing Wang

机构 * School of Software Engineering, Huazhong University of Science and Technology(华中科技大学软件学院) School of Cyber Science and Engineering, Huazhong University of Science and Technology(华中科技大学网络安全学院) College of Computing and Data Science, Nanyang Technological University(南洋理工大学计算与数据科学学院) Faculty of Data Science, City University of Macau(澳门城市大学数据科学学院)

专题命中 提示注入 :prompt injection(title,abstract);分类 cs.AI

Comments Accepted to EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 幻觉与事实性 1 篇

2509.07475 2025-09-10 cs.CL cs.AI 62%

HALT-RAG: A Task-Adaptable Framework for Hallucination Detection with Calibrated NLI Ensembles and Abstention

Saumya Goswami, Siddharth Kurra

专题命中 幻觉与事实性 :safety(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 安全评测 13 篇

2508.17450 2025-09-10 cs.CL cs.CY 84%

Persuasion Dynamics in LLMs: Investigating Robustness and Adaptability in Knowledge and Safety with DuET-PD

Bryan Chen Zhengyu Tan, Daniel Wai Kit Chin, Zhengyuan Liu, Nancy F. Chen, Roy Ka-Wei Lee

机构 * Singapore University of Technology and Design (SUTD)(新加坡科技设计大学) Institute for Infocomm Research (I2R), A*STAR, Singapore(信息通信研究院)

专题命中 安全评测 :safety(title,abstract);DPO(abstract);分类 cs.CL、cs.CY

Comments To appear at EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.07315 2025-09-10 cs.CR cs.SE 82%

SafeToolBench: Pioneering a Prospective Benchmark to Evaluating Tool Utilization Safety in LLMs

Hongfei Xia, Hongru Wang, Zeming Liu, Qian Yu, Yuhang Guo, Haifeng Wang

专题命中 安全评测 :safety(title,abstract);trustworthy(abstract)

Comments 18 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00700 2025-09-10 cs.CV 78%

Prompt the Unseen: Evaluating Visual-Language Alignment Beyond Supervision

Raehyuk Jung, Seungjun Yu, Hyunjung Shim

专题命中 安全评测 :alignment(title,abstract)

Comments Link to publicly available codes is added

详情

展开后加载摘要…

URL PDF HTML 收藏
2101.12100 2025-09-10 cs.LG cs.AI 73%

Increasing the Confidence of Deep Neural Networks by Coverage Analysis

Giulio Rossolini, Alessandro Biondi, Giorgio Buttazzo

专题命中 安全评测 :safety(abstract);trustworthy(abstract);分类 cs.AI、cs.LG

Journal ref IEEE Transactions on Software Engineering ( Volume: 49, Issue: 2, 01 February 2023)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.18933 2025-09-10 cs.AI cs.CR cs.CY cs.LG 67%

VISION: Robust and Interpretable Code Vulnerability Detection Leveraging Counterfactual Augmentation

David Egea, Barproda Halder, Sanghamitra Dutta

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.CY、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.07017 2025-09-10 cs.AI cs.CL cs.LG 67%

From Eigenmodes to Proofs: Integrating Graph Spectral Operators with Symbolic Interpretable Reasoning

Andrew Kiruluta, Priscilla Burity

机构 * Andrew Kiruluta and Priscilla Burity(独立研究者)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.07473 2025-09-10 cs.AI 57%

SheetDesigner: MLLM-Powered Spreadsheet Layout Generation with Rule-Based and Vision-Based Reflection

Qin Chen, Yuanyi Ren, Xiaojun Ma, Mugeng Liu, Han Shi, Dongmei Zhang

机构 * Peking University(北京大学) Microsoft(微软)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

Comments Accepted to EMNLP 2025 Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.07127 2025-09-10 cs.GR cs.AI cs.CV 57%

SVGauge: Towards Human-Aligned Evaluation for SVG Generation

Leonardo Zini, Elia Frigieri, Sebastiano Aloscari, Marcello Generali, Lorenzo Dodi, Robert Dosen, Lorenzo Baraldi

机构 * University of Modena and Reggio Emilia(摩德纳和雷吉奥艾米利亚大学) Doxee S.p.A.(Doxee公司)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

Comments Accepted at 23rd edition of International Conference on Image Analysis and Processing 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05702 2025-09-10 cs.MA cs.AI cs.SY eess.SY 57%

Grid-Agent: An LLM-Powered Multi-Agent System for Power Grid Control

Yan Zhang, Ahmad Mohammad Saber, Amr Youssef, Deepa Kundur

机构 * Department of Electrical and Computer Engineering, University of Toronto(电气与计算机工程系,多伦多大学) Concordia Institute for Information Systems Engineering(康卡迪亚信息系统工程研究所) Concordia University(康卡迪亚大学)

专题命中 安全评测 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.07039 2025-09-10 cs.LG cs.CV 57%

Benchmarking Vision Transformers and CNNs for Thermal Photovoltaic Fault Detection with Explainable AI Validation

Serra Aksoy

专题命中 安全评测 :alignment(abstract);分类 cs.LG

Comments 28 Pages, 4 Figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.07026 2025-09-10 cs.LO cs.AI 57%

Contradictions

Yang Xu, Shuwei Chen, Xiaomei Zhong, Jun Liu, Xingxing He

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

Comments 37 Pages,9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.07010 2025-09-10 cs.CV cs.AI cs.ET 57%

Human-in-the-Loop: Quantitative Evaluation of 3D Models Generation by Large Language Models

Ahmed R. Sadik, Mariusz Bujny

机构 * Honda Research Institute Europe - Germany(本田欧洲研究机构)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.07323 2025-09-10 cs.SD cs.CR 50%

When Fine-Tuning is Not Enough: Lessons from HSAD on Hybrid and Adversarial Audio Spoof Detection

Bin Hu, Kunyang Huang, Daehan Kwak, Meng Xu, Kuan Huang

机构 * Department of Computer Science and Technology, Kean University, USA(计算机科学与技术系,凯恩大学,美国) Department of Computer Science and Technology, Wenzhou-Kean University, China(计算机科学与技术系,温州-凯恩大学,中国)

专题命中 安全评测 :trustworthy(abstract)

Comments 13 pages, 11 figures.This work has been submitted to the IEEE for possible publication

详情

展开后加载摘要…

URL PDF HTML 收藏

6. AI治理与伦理 2 篇

2509.07022 2025-09-10 cs.CY cs.AI 81%

Preventing Another Tessa: Modular Safety Middleware For Health-Adjacent AI Assistants

Pavan Reddy, Nithin Reddy

专题命中 AI治理与伦理 :safety(title,abstract);分类 cs.AI、cs.CY

Comments 7 pages content, 1 page reference, 1 figure, Accepted at AAAI Fall Symposium Series

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.07006 2025-09-10 cs.CY cs.AI cs.CL cs.LG 78%

ArGen: Auto-Regulation of Generative AI via GRPO and Policy-as-Code

Kapil Madan

机构 * Principled Evolution(原则进化)

专题命中 AI治理与伦理 :alignment(abstract,comments);safety(abstract);分类 cs.CL、cs.AI、cs.CY

Comments 53 pages, 7 figures, 8 tables. Open-source implementation available at: https://github.com/Principled-Evolution/argen-demo. Work explores the integration of policy-as-code for AI alignment, with a case study in culturally-nuanced, ethical AI using Dharmic principles

详情

展开后加载摘要…

URL PDF HTML 收藏

7. 其他安全 10 篇

2509.06982 2025-09-10 cs.LG cs.AI cs.CL 89%

CARE: Decoding Time Safety Alignment via Rollback and Introspection Intervention

Xiaomeng Hu, Fei Huang, Chenhan Yuan, Junyang Lin, Tsung-Yi Ho

机构 * Alibaba Group(阿里巴巴集团) The Chinese University of Hong Kong(香港中文大学)

专题命中 其他安全 :alignment(title,abstract);safety(title,abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.06966 2025-09-10 eess.SP cs.AI cs.LG 82%

Cross-device Zero-shot Label Transfer via Alignment of Time Series Foundation Model Embeddings

Neal G. Ravindra, Arijit Sehanobish

机构 * Independent Researcher(独立研究者)

专题命中 其他安全 :alignment(title,abstract);分类 cs.AI、cs.LG

Comments 5 pages, 3 figures, 1 table. tl;dr: Adversarial alignment of Time-Series Foundation Model (TSFM) embeddings enables transfer of high-quality clinical labels from medical-grade to consumer-grade wearables, enabling zero-shot prediction of gestational age without requiring paired data

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.07588 2025-09-10 cs.CL cs.AI 81%

BALI: Enhancing Biomedical Language Representations through Knowledge Graph and Language Model Alignment

Andrey Sakhovskiy, Elena Tutubalina

机构 * AIRI Sber AI ISP RAS Research Center for Trusted AI(俄罗斯科学院信息与系统研究所可信人工智能研究中心)

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI

Comments 9 pages, 1 figure, published in "The 48th International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR 2025)"

Journal ref Proceedings of the 48th International ACM SIGIR Conference on Research and Development in Information Retrieval (2025). Association for Computing Machinery, 1152-1164

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.07642 2025-09-10 cs.AI 79%

Getting In Contract with Large Language Models -- An Agency Theory Perspective On Large Language Model Alignment

Sascha Kaltenpoth, Oliver Müller

机构 * Paderborn University, Department of Business Administration and Economics(帕德博恩大学商业管理与经济学系)

专题命中 其他安全 :alignment(title,abstract);分类 cs.AI

Comments Presented at the 19th International Conference on Wirtschaftsinformatik 2024, Würzburg, Germany https://aisel.aisnet.org/wi2024/91/

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.07311 2025-09-10 cs.CL cs.AI 62%

Does This Look Familiar to You? Knowledge Analysis via Model Internal Representations

Sihyun Park

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.02920 2025-09-10 cs.LG cs.CY cs.ET cs.SY eess.SY 62%

Event Detection and Classification for Long Range Sensing of Elephants Using Seismic Signal

Jaliya L. Wijayaraja, Janaka L. Wijekoon, Malitha Wijesundara

机构 * Sri Lanka Institute of Information Technology(斯里兰卡信息技术研究所) Victorian Institute of Technology(维多利亚技术学院) Department of System Design Engineering, Keio University(系统设计工程系,庆应大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CY、cs.LG

Comments This article has been accepted for publication in IEEE Access

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.07846 2025-09-10 cs.AI 57%

Aligning LLMs for the Classroom with Knowledge-Based Retrieval -- A Comparative RAG Study

Amay Jain, Liu Cui, Si Chen

机构 * Student(学生) Downingtown STEM Academy Department of Computer Science(计算机科学系) West Chester University of Pennsylvania(宾夕法尼亚州韦斯特切斯特大学)

专题命中 其他安全 :alignment(abstract);分类 cs.AI

Comments This work has been submitted to the IEEE for possible publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.07334 2025-09-10 cs.HC 50%

SpecifyUI: Supporting Iterative UI Design Intent Expression through Structured Specifications and Generative AI

Yunnong Chen, Chengwei Shi, Liuqing Chen

专题命中 其他安全 :alignment(abstract)

Comments 27 pages, 12 figures

详情

展开后加载摘要…

URL PDF HTML 收藏