arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-08-11 至 2025-08-11 共收录 12 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 12 篇

2508.06124 2025-08-11 cs.CL 83%

AURA: Affordance-Understanding and Risk-aware Alignment Technique for Large Language Models

Sayantan Adak, Pratyush Chatterjee, Somnath Banerjee, Rima Hazra, Somak Aditya, Animesh Mukherjee

专题命中 安全评测 :alignment(title,abstract);safety(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05855 2025-08-11 cs.AI cs.RO 79%

Safety of Embodied Navigation: A Survey

Zixia Wang, Jia Hu, Ronghui Mu

机构 * University of Exeter(埃克塞特大学)

专题命中 安全评测 :safety(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.03834 2025-08-11 cs.RO cs.CV 78%

CARE: Enhancing Safety of Visual Navigation through Collision Avoidance via Repulsive Estimation

Joonkyung Kim, Joonyeol Sim, Woojun Kim, Katia Sycara, Changjoo Nam

机构 * Department of Electronic Engineering, Sogang University(电子工程系,首尔大学) Robotics Institute, Carnegie Mellon University(机器人研究所,卡内基梅隆大学)

专题命中 安全评测 :safety(title,abstract)

Comments 16 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.00381 2025-08-11 cs.CV cs.AI cs.CE cs.LG 73%

Advancing Welding Defect Detection in Maritime Operations via Adapt-WeldNet and Defect Detection Interpretability Analysis

Kamal Basha S, Athira Nambiar

机构 * Department of Computational Intelligence, Faculty of Engineering and Technology, SRM Institute of Science and Technology(计算智能系,工程与技术学院,SRM科学与技术学院)

专题命中 安全评测 :safety(abstract);trustworthy(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.13677 2025-08-11 cs.CL cs.AI cs.LG 67%

Revisiting Uncertainty Quantification Evaluation in Language Models: Spurious Interactions with Response Length Bias Results

Andrea Santilli, Adam Golinski, Michael Kirchhof, Federico Danieli, Arno Blaas, Miao Xiong, Luca Zappella, Sinead Williamson

专题命中 安全评测 :safety(abstract);分类 cs.CL、cs.AI、cs.LG

Comments Accepted at ACL 2025 (Main)

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.14119 2025-08-11 cs.CL cs.AI 62%

Autonomous Structural Memory Manipulation for Large Language Models Using Hierarchical Embedding Augmentation

Derek Yotheringhay, Alistair Kirkland, Humphrey Kirkbride, Josiah Whitesteeple

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

Comments arXiv admin note: This paper has been withdrawn by arXiv due to disputed and unverifiable authorship

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.11417 2025-08-11 cs.CL cs.AI 62%

Neural Contextual Reinforcement Framework for Logical Structure Language Generation

Marcus Irvin, William Cooper, Edward Hughes, Jessica Morgan, Christopher Hamilton

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

Comments arXiv admin note: This paper has been withdrawn by arXiv due to disputed and unverifiable authorship

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06155 2025-08-11 cs.CL 57%

Semantic and Structural Analysis of Implicit Biases in Large Language Models: An Interpretable Approach

Renhan Zhang, Lian Lian, Zhen Qi, Guiran Liu

专题命中 安全评测 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05987 2025-08-11 cs.CL 57%

Adversarial Topic-aware Prompt-tuning for Cross-topic Automated Essay Scoring

Chunyun Zhang, Hongyan Zhao, Chaoran Cui, Qilong Song, Zhiqing Lu, Shuai Gong, Kailin Liu

机构 * Shandong University of Finance and Economics(山东财经大学) University of Toronto(多伦多大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.08947 2025-08-11 cs.CL 57%

Structured Convergence in Large Language Model Representations via Hierarchical Latent Space Folding

Fenella Harcourt, Naderdel Piero, Gilbert Sutherland, Daphne Holloway, Harriet Bracknell, Julian Ormsby

专题命中 安全评测 :alignment(abstract);分类 cs.CL

Comments arXiv admin note: This paper has been withdrawn by arXiv due to disputed and unverifiable authorship

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05687 2025-08-11 cs.MA cs.AI 57%

Risk Analysis Techniques for Governed LLM-based Multi-Agent Systems

Alistair Reid, Simon O'Callaghan, Liam Carroll, Tiberio Caetano

机构 * Gradient Institute Ltd.(梯度研究所有限公司)

专题命中 安全评测 :red teaming(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06152 2025-08-11 cs.CV 50%

VISTAR:A User-Centric and Role-Driven Benchmark for Text-to-Image Evaluation

Kaiyuan Jiang, Ruoxi Sun, Ying Cao, Yuqi Xu, Xinran Zhang, Junyan Guo, ChengSheng Deng

机构 * Peking University(北京大学) LinkSure University of Glasgow(格拉斯哥大学) Boston University(波士顿大学)

专题命中 安全评测 :alignment(abstract)

Comments 17 pages,8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏