arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-09-17 至 2025-09-17 共收录 11 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 11 篇

2509.12936 2025-09-17 cs.LG cs.CL 90%

Rethinking the Evaluation of Alignment Methods: Insights into Diversity, Generalisation, and Safety

Denis Janiak, Julia Moska, Dawid Motyka, Karolina Seweryn, Paweł Walkowiak, Bartosz Żuk, Arkadiusz Janz

机构 * Wroclaw University of Science and Technology (WUST)(沃拉布大学科学与技术学院) National Research Institute (NASK)(国家研究 institute) Institute of Computer Science, Polish Academy of Sciences (IPI PAN)(波兰科学院计算机科学研究所)

专题命中 安全评测 :alignment(title,abstract);safety(title,abstract);DPO(abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.13244 2025-09-17 cs.CL 79%

Evaluating LLM Alignment on Personality Inference from Real-World Interview Data

Jianfeng Zhu, Julina Maharjan, Xinyu Li, Karin G. Coifman, Ruoming Jin

专题命中 安全评测 :alignment(title,abstract);分类 cs.CL

Comments 8 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.12233 2025-09-17 cs.CR cs.AI cs.ET cs.LG cs.NI 76%

Towards Trustworthy Agentic IoEV: AI Agents for Explainable Cyberthreat Mitigation and State Analytics

Meryem Malak Dif, Mouhamed Amine Bouchiha, Abdelaziz Amara Korba, Yacine Ghamri-Doudane

机构 * L3i - La Rochelle University, La Rochelle, France(L3i - 拉罗谢尔大学)

专题命中 安全评测 :trustworthy(title);分类 cs.AI、cs.LG

Comments 10 pages, 7 figures, Accepted at LCN'25

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.12740 2025-09-17 cs.RO cs.AI cs.ET cs.LG cs.SY eess.SY 62%

Deep Generative and Discriminative Digital Twin endowed with Variational Autoencoder for Unsupervised Predictive Thermal Condition Monitoring of Physical Robots in Industry 6.0 and Society 6.0

Eric Guiffo Kaigom

机构 * Department of Computer Science \& Engineering, Frankfurt University of Applied Sciences, Frankfurt a.M., Germany (e-mail: ).

专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.LG

Comments $©$ 2025 the authors. This work has been accepted to the to the 10th IFAC Symposium on Mechatronic Systems & 14th IFAC Symposium on Robotics July 15-18, 2025 || Paris, France for publication under a Creative Commons Licence CC-BY-NC-ND

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.12259 2025-09-17 cs.LG cs.AI quant-ph 62%

Quantum-Inspired Stacked Integrated Concept Graph Model (QISICGM) for Diabetes Risk Prediction

Kenneth G. Young

机构 * II (September 12, 2025)(II)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG

Comments 13 pages, 3 figures, includes performance tables and visualizations. Proposes a Quantum-Inspired Stacked Integrated Concept Graph Model (QISICGM) that integrates phase feature mapping, self-improving concept graphs, and neighborhood sequence modeling within a stacked ensemble. Demonstrates improved F1 and AUC on an augmented PIMA Diabetes dataset with efficient CPU inference

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.12612 2025-09-17 cs.AI 57%

GBV-SQL: Guided Generation and SQL2Text Back-Translation Validation for Multi-Agent Text2SQL

Daojun Chen, Xi Wang, Shenyuan Ren, Qingzhi Ma, Pengpeng Zhao, An Liu

专题命中 安全评测 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.12459 2025-09-17 cs.CL 57%

Does Language Model Understand Language?

Suvojit Acharjee, Utathya Aich, Asfak Ali

机构 * Institute of Engineering and Management(工程与管理研究所) Jadavpur University(贾达沃大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.10483 2025-09-17 cs.SE cs.AI cs.PL 57%

Enhancing Automated Loop Invariant Generation for Complex Programs with Large Language Models

Ruibang Liu, Minyu Chen, Ling-I Wu, Jingyu Ke, Guoqiang Li

机构 * Shanghai Jiao Tong University(上海交通大学)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

Comments 26 pages, 11 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.12912 2025-09-17 cs.RO 50%

Spotting the Unfriendly Robot -- Towards better Metrics for Interactions

Raphael Wenzel, Malte Probst

机构 * Honda Research Institute Europe GmbH(本田欧洲研究院)

专题命中 安全评测 :safety(abstract)

Comments Presented at 2025 IEEE Conference on Robotics and Automation (ICRA) Workshop: Advances in Social Navigation: Planning, HRI and Beyond

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.12492 2025-09-17 cs.CV 50%

Evaluating Robustness of Vision-Language Models Under Noisy Conditions

Purushoth, Alireza

机构 * University of Nevada Reno(内华达大学拉斯维加斯分校)

专题命中 安全评测 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11292 2025-09-17 cs.CV 50%

Leveraging Geometric Priors for Unaligned Scene Change Detection

Ziling Liu, Ziwei Chen, Mingqi Gao, Jinyu Yang, Feng Zheng

机构 * Southern University of Science and Technology(南方科技大学) University of Sheffield(谢菲尔德大学) Spatialtemporal AI(时空AI)

专题命中 安全评测 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏