arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-09-08 至 2025-09-08 共收录 45 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 17 篇

2509.04499 2025-09-08 cs.CL cs.AI 62%

DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence

Pranav Narayanan Venkit, Philippe Laban, Yilun Zhou, Kung-Hsiang Huang, Yixin Mao, Chien-Sheng Wu

机构 * Salesforce AI Research(Salesforce AI研究)

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL、cs.AI

Comments arXiv admin note: text overlap with arXiv:2410.22349

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04794 2025-09-08 cs.CL 57%

Personality as a Probe for LLM Evaluation: Method Trade-offs and Downstream Effects

Gunmay Handa, Zekun Wu, Adriano Koshiyama, Philip Treleaven

机构 * University College London(伦敦大学学院)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04619 2025-09-08 eess.SY cs.SY math.DS 50%

$\mathcal{L}_1$-DRAC: Distributionally Robust Adaptive Control

Aditya Gahlawat, Sambhu H. Karumanchi, Naira Hovakimyan

专题命中 安全评测 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

2. AI治理与伦理 2 篇

2509.04993 2025-09-08 cs.MA cs.AI 57%

LLM Enabled Multi-Agent System for 6G Networks: Framework and Method of Dual-Loop Edge-Terminal Collaboration

Zheyan Qu, Wenbo Wang, Zitong Yu, Boquan Sun, Yang Li, Xing Zhang

机构 * Beijing University of Posts and Telecommunications(北京邮电大学)

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI

Comments This paper has been accepted by IEEE Communications Magazine

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20201 2025-09-08 cs.CL 57%

Social Bias in Multilingual Language Models: A Survey

Lance Calvin Lim Gamboa, Yue Feng, Mark Lee

机构 * School of Computer Science, University of Birmingham(伯明翰大学计算机科学学院) Department of Information Systems and Computer Science, Ateneo de Manila University(马尼拉大学信息系统与计算机科学系)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL

Comments Accepted into EMNLP 2025 Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 其他安全 10 篇

2408.11813 2025-09-08 cs.CV 78%

SEA: Supervised Embedding Alignment for Token-Level Visual-Textual Integration in MLLMs

Yuanyang Yin, Yaqi Zhao, Yajie Zhang, Yuanxing Zhang, Ke Lin, Jiahao Wang, Xin Tao, Pengfei Wan, Wentao Zhang, Feng Zhao

机构 * University of Science and Technology of China(中国科学技术大学) Peking University(北京大学) Kuaishou Technology(快手科技)

专题命中 其他安全 :alignment(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04876 2025-09-08 cs.AI 74%

OSC: Cognitive Orchestration through Dynamic Knowledge Alignment in Multi-Agent LLM Collaboration

Jusheng Zhang, Yijia Fan, Kaitong Cai, Xiaofei Sun, Keze Wang

机构 * Sun Yat-sen University(中山大学) Alibaba Group(阿里巴巴集团)

专题命中 其他安全 :alignment(title);分类 cs.AI

Comments Accepted at EMNLP 2025 (Long Paper)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04500 2025-09-08 cs.CL cs.AI 62%

Context Engineering for Trustworthiness: Rescorla Wagner Steering Under Mixed and Inappropriate Contexts

Rushi Wang, Jiateng Liu, Cheng Qian, Yifan Shen, Yanzhou Pan, Zhaozhuo Xu, Ahmed Abbasi, Heng Ji, Denghui Zhang

专题命中 其他安全 :safety(abstract);分类 cs.CL、cs.AI

Comments 36 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.05051 2025-09-08 quant-ph cs.LG 57%

QCA-MolGAN: Quantum Circuit Associative Molecular GAN with Multi-Agent Reinforcement Learning

Aaron Mark Thomas, Yu-Cheng Chen, Hubert Okadome Valencia, Sharu Theresa Jose, Ronin Wu

机构 * Department of Computer Science(计算机科学系) University of Birmingham(伯明翰大学) QunaSys Europe(QunaSys欧洲分公司) Hon Hai Research Institute(鸿海研究机构) Departement of Computer Science(计算机科学系)

专题命中 其他安全 :alignment(abstract);分类 cs.LG

Comments Accepted to the proceedings of IEEE Quantum Artificial Intelligence, 6 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04942 2025-09-08 cs.LG 57%

Ontology-Aligned Embeddings for Data-Driven Labour Market Analytics

Heinke Hihn, Dennis A. V. Dittrich, Carl Jeske, Cayo Costa Sobral, Helio Pais, Timm Lochmann

机构 * IU International University of Applied Sciences, Department of Computer Science and Engineering(国际应用科学大学计算机科学与工程系) The Stepstone Group, AI Labs(Stepstone集团人工智能实验室)

专题命中 其他安全 :alignment(abstract);分类 cs.LG

Comments Workshop SIG Knowledge Management (FG WM) at KI2025, Potsdam, Germany

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04751 2025-09-08 cs.IR cs.LG 57%

Multimodal Foundation Model-Driven User Interest Modeling and Behavior Analysis on Short Video Platforms

Yushang Zhao, Yike Peng, Li Zhang, Qianyi Sun, Zhihui Zhang, Yingying Zhuang

专题命中 其他安全 :alignment(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.07173 2025-09-08 cs.AI 57%

Translating Federated Learning Algorithms in Python into CSP Processes Using ChatGPT

Miroslav Popovic, Marko Popovic, Miodrag Djukic, Ilija Basicevic

专题命中 其他安全 :safety(abstract);分类 cs.AI

Comments 6 pages, 4 tables; Published by IEEE Xplore

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.05116 2025-09-08 cs.ET cs.RO 50%

Analyzing Gait Adaptation with Hemiplegia Simulation Suits and Digital Twins

Jialin Chen, Jeremie Clos, Dominic Price, Praminda Caleb-Solly

机构 * University of Nottingham, UK(诺丁汉大学)

专题命中 其他安全 :safety(abstract)

Comments 7 pages, accepted at EMBC 2025, presented at the conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.05004 2025-09-08 cs.CV 50%

Interpretable Deep Transfer Learning for Breast Ultrasound Cancer Detection: A Multi-Dataset Study

Mohammad Abbadi, Yassine Himeur, Shadi Atalla, Wathiq Mansoor

机构 * College of Engineering and Information Technology(工程与信息技术学院)

专题命中 其他安全 :safety(abstract)

Comments 6 pages, 2 figures and 1 table

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04957 2025-09-08 cs.CV cs.MM cs.SD eess.AS 50%

Efficient Video-to-Audio Generation via Multiple Foundation Models Mapper

Gehui Chen, Guan'an Wang, Xiaowen Huang, Jitao Sang

机构 * School of Computer Science Technology, Beijing Jiaotong University Beijing China Beijing Key Laboratory of Traffic Data Mining Key Laboratory of Big Data \& Artificial Intelligence in Transportation, Ministry of Education Beijing China Technology, Beijing Jiaotong University Key Laboratory of Big Data \& Artificial Intelligence in Transportation, Ministry of Education

专题命中 其他安全 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏