arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-11-06 至 2025-11-06 共收录 7 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 7 篇

2211.01446 2025-11-06 cs.LG 74%

Trustworthy Representation Learning via Information Funnels and Bottlenecks

João Machado de Freitas, Bernhard C. Geiger

机构 * Christian Doppler Laboratory for Dependable Intelligent Systems in Harsh Environments(可信智能系统在恶劣环境中的克里斯蒂安·多普勒实验室) Graz University of Technology(格拉茨技术大学) Know Center Research GmbH(Know Center研究有限责任公司)

专题命中 安全评测 :trustworthy(title);分类 cs.LG

Comments Published in Machine Learning (Springer), vol. 114, no. 12, Article 267, 2025

Journal ref Mach Learn 114, 267 (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.03641 2025-11-06 cs.CR cs.AI cs.CL cs.CY 67%

Watermarking Large Language Models in Europe: Interpreting the AI Act in Light of Technology

Thomas Souverain

机构 * Department of AI Ethics, CEA Paris-Saclay(人工智能伦理系,CEA巴黎萨克雷)

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL、cs.AI、cs.CY

Comments 17 pages, 2 Tables and 2 Pictures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21861 2025-11-06 cs.LG cs.AI cs.CL 67%

The Mirror Loop: Recursive Non-Convergence in Generative Reasoning Systems

Bentley DeVilling

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments 18 pages, 2 figures. Category: cs.LG. Code and data: https://github.com/Course-Correct-Labs/mirror-loop

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.03051 2025-11-06 cs.AI cs.IR 57%

No-Human in the Loop: Agentic Evaluation at Scale for Recommendation

Tao Zhang, Kehui Yao, Luyi Ma, Jiao Chen, Reza Yousefi Maragheh, Kai Zhao, Jianpeng Xu, Evren Korpeoglu, Sushant Kumar, Kannan Achan

机构 * Walmart Global Tech(沃尔玛全球技术)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

Comments 4 page, NeurIPS 2025 Workshop: Evaluating the Evolving LLM Lifecycle

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.20462 2025-11-06 cs.AI 57%

TAMO: Fine-Grained Root Cause Analysis via Tool-Assisted LLM Agent with Multi-Modality Observation Data in Cloud-Native Systems

Xiao Zhang, Qi Wang, Mingyi Li, Yuan Yuan, Mengbai Xiao, Fuzhen Zhuang, Dongxiao Yu

机构 * School of Computer Science and Technology, Shandong University(山东大学计算机科学与技术学院) Institute of Artificial Intelligence, Beihang University(北京航空航天大学人工智能研究院)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.07955 2025-11-06 cs.HC 50%

Implementation Considerations for Automated AI Grading of Student Work

Zewei Tian, Alex Liu, Lief Esbenshade, Shawon Sarkar, Zachary Zhang, Kevin He, Min Sun

专题命中 安全评测 :trustworthy(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.02996 2025-11-06 cs.CV 50%

SCALE-VLP: Soft-Weighted Contrastive Volumetric Vision-Language Pre-training with Spatial-Knowledge Semantics

Ailar Mahdizadeh, Puria Azadi Moghadam, Xiangteng He, Shahriar Mirabbasi, Panos Nasiopoulos, Leonid Sigal

机构 * University of British Columbia(不列颠哥伦比亚大学) Vector Institute for AI(人工智能向量研究所)

专题命中 安全评测 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏