arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-10-28 至 2025-10-28 共收录 84 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 25 篇

2505.12116 2025-10-28 cs.CL 57%

A Multi-Task Benchmark for Abusive Language Detection in Low-Resource Settings

Fitsum Gaim, Hoyun Song, Huije Lee, Changgeon Ko, Eui Jun Hwang, Jong C. Park

机构 * Korea Advanced Institute of Science and Technology (KAIST)(韩国科学技术院)

专题命中 安全评测 :safety(abstract);分类 cs.CL

Comments Accepted at NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21722 2025-10-28 cs.HC cs.AI 57%

AquaVLM: Improving Underwater Situation Awareness with Mobile Vision Language Models

Beitong Tian, Lingzhi Zhao, Bo Chen, Haozhen Zheng, Jingcheng Yang, Mingyuan Wu, Deepak Vasisht, Klara Nahrstedt

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

专题命中 安全评测 :safety(abstract);分类 cs.AI

Comments 12 pages, 10 figures, under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.09110 2025-10-28 cs.CV 56%

Pulling Back the Curtain: Unsupervised Adversarial Detection via Contrastive Auxiliary Networks

Eylon Mizrahi, Raz Lapid, Moshe Sipper

机构 * Ben-Gurion University(本·古里安大学) DeepKeep

专题命中 安全评测 :safety(abstract);trustworthy(journal_ref)

Comments Accepted for Oral Presentation at SafeMM-AI @ ICCV 2025 (Spotlight)

Journal ref ICCV 2025 Workshop on SafeMM-AI: Safe and Trustworthy Multimodal AI Systems

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23032 2025-10-28 cs.CE 50%

P1GPT: a multi-agent LLM workflow module for multi-modal financial information analysis

Chen-Che Lu, Yun-Cheng Chou, Teng-Ruei Chen

专题命中 安全评测 :trustworthy(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22800 2025-10-28 cs.NE 50%

Probing the Representational Geometry of Color Qualia: Dissociating Pure Perception from Task Demands in Brains and AI Models

Jing Xu

专题命中 安全评测 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22785 2025-10-28 cs.CV 50%

Self-Calibrated Consistency can Fight Back for Adversarial Robustness in Vision-Language Models

Jiaxiang Liu, Jiawei Du, Xiao Liu, Prayag Tiwari, Mingkun Xu

机构 * Guangdong Institute of Intelligence Science and Technology(广东智能科学与技术研究院) Agency for Science, Technology and Research(科技研究局) School of Information Technology(信息技术学院)

专题命中 安全评测 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.20625 2025-10-28 cs.CV 50%

T2ICount: Enhancing Cross-modal Understanding for Zero-Shot Counting

Yifei Qian, Zhongliang Guo, Bowen Deng, Chun Tong Lei, Shuai Zhao, Chun Pong Lau, Xiaopeng Hong, Michael P. Pound

机构 * University of Nottingham(诺丁汉大学) University of St Andrews(圣安德鲁大学) City University of Hong Kong(香港城市大学) Nanyang Technology University(南洋理工大学) Harbin Institute of Technology(哈尔滨工业大学)

专题命中 安全评测 :alignment(abstract)

Comments Accepted by CVPR2025

详情

展开后加载摘要…

URL PDF HTML 收藏

2. AI治理与伦理 5 篇

2506.10192 2025-10-28 cs.AI 83%

Towards Responsible AI: Advances in Safety, Fairness, and Accountability of Autonomous Systems

Filip Cano

专题命中 AI治理与伦理 :safety(title,abstract);trustworthy(abstract);分类 cs.AI

Comments 204 pages, 38 figures, PhD Thesis

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23245 2025-10-28 cs.HC cs.MA 71%

Multi-Stakeholder Alignment in LLM-Powered Collaborative AI Systems: A Multi-Agent Framework for Intelligent Tutoring

Alexandre P Uchoa, Carlo E T Oliveira, Claudia L R Motta, Daniel Schneider

专题命中 AI治理与伦理 :alignment(title)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22823 2025-10-28 cs.CL cs.AI 62%

Cross-Lingual Stability and Bias in Instruction-Tuned Language Models for Humanitarian NLP

Poli Nemkova, Amrit Adhikari, Matthew Pearson, Vamsi Krishna Sadu, Mark V. Albert

机构 * University of North Texas(北卡罗来纳大学达顿分校) Davidson College(戴维森学院)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22286 2025-10-28 cs.CY 57%

Hybrid Instructor Ai Assessment In Academic Projects: Efficiency, Equity, And Methodological Lessons

Hugo Roger Paz

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CY

Comments 12 pages, in Spanish language, 0 figures, 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21967 2025-10-28 cs.HC cs.CY 57%

We Need Accountability in Human-AI Agent Relationships

Benjamin Lange, Geoff Keeling, Arianna Manzini, Amanda McCroskery

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 其他安全 12 篇

2402.01342 2025-10-28 cs.LG stat.ML 79%

Improving Model Fusion by Training-time Neuron Alignment with Fixed Neuron Anchors

Zexi Li, Zhiqi Li, Jie Lin, Tao Shen, Jun Xiao, Yike Guo, Tao Lin, Chao Wu

机构 * Zhejiang University(浙江大学) Georgia Institute of Technology(佐治亚理工学院) Westlake University(西湖大学) Hong Kong University of Science and Technology(香港科技大学)

专题命中 其他安全 :alignment(title,abstract);分类 cs.LG

Comments IEEE Transactions on Pattern Analysis and Machine Intelligence

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.15648 2025-10-28 cs.CR cs.LG cs.SE 79%

deepSURF: Detecting Memory Safety Vulnerabilities in Rust Through Fuzzing LLM-Augmented Harnesses

Georgios Androutsopoulos, Antonio Bianchi

机构 * Purdue University(普渡大学)

专题命中 其他安全 :safety(title,abstract);分类 cs.LG

Comments At IEEE S&P 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.19586 2025-10-28 cs.CL q-bio.NC 76%

Distinct social-linguistic processing between humans and large audio-language models: Evidence from model-brain alignment

Hanlin Wu, Xufeng Duan, Zhenguang Cai

机构 * Department of Linguistics and Modern Languages, The Chinese University of Hong Kong(语言学与现代语言系,香港中文大学) Brain and Mind Institute, The Chinese University of Hong Kong(脑与心智研究所,香港中文大学)

专题命中 其他安全 :alignment(title,comments);分类 cs.CL

Comments Hanlin Wu, Xufeng Duan, and Zhenguang Cai. 2025. Distinct social-linguistic processing between humans and large audio-language models: Evidence from model-brain alignment. In Proceedings of the Workshop on Cognitive Modeling and Computational Linguistics, pages 135-143, Albuquerque, New Mexico, USA. Association for Computational Linguistics. https://aclanthology.org/2025.cmcl-1.18/

Journal ref In Proceedings of CMCL, pages 135-143, ACL (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23160 2025-10-28 cs.CL 57%

ENTP: Enhancing Low-Quality SFT Data via Neural-Symbolic Text Purge-Mix

Zile Yang, Ling Li, Na Di, Jinlong Pang, Yao Zhou, Hao Cheng, Bo Han, Jiaheng Wei

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州)) University of California, Santa Cruz(加州大学圣克鲁兹分校) Hong Kong Baptist University(香港 Baptist 大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23988 2025-10-28 cs.AI cs.DB 57%

LLM/Agent-as-Data-Analyst: A Survey

Zirui Tang, Weizheng Wang, Zihang Zhou, Yang Jiao, Bangrui Xu, Boyu Niu, Dayou Zhou, Xuanhe Zhou, Guoliang Li, Yeye He, Wei Zhou, Yitong Song, Cheng Tan, Xue Yang, Chunwei Liu, Bin Wang, Conghui He, Xiaoyang Wang, Fan Wu

机构 * Shanghai Jiao Tong University(上海交通大学) Tsinghua University(清华大学) Microsoft Research(微软研究院) MIT CSAIL(麻省理工学院计算机科学与人工智能实验室) Shanghai AI Laboratory(上海人工智能实验室) Fudan University(复旦大学)

专题命中 其他安全 :alignment(abstract);分类 cs.AI

Comments 31 page, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.20014 2025-10-28 cs.CL 57%

Does Rationale Quality Matter? Enhancing Mental Disorder Detection via Selective Reasoning Distillation

Hoyun Song, Huije Lee, Jisu Shin, Sukmin Cho, Changgeon Ko, Jong C. Park

机构 * School of Computing(计算学院) Korea Advanced Institute of Science and Technology(韩国科学技术院)

专题命中 其他安全 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.15742 2025-10-28 cs.LG cs.AR eess.SP 57%

DeepVigor+: Scalable and Accurate Semi-Analytical Fault Resilience Analysis for Deep Neural Network

Mohammad Hasan Ahmadilivani, Jaan Raik, Masoud Daneshtalab, Maksim Jenihhin

机构 * Tallinn University of Technology(塔林技术大学) Mälardalen University(马尔默大学)

专题命中 其他安全 :safety(abstract);分类 cs.LG

Comments 14 pages, 9 figures, 8 tables, 16 equations. The source code is accessible via: https://github.com/mhahmadilivany/DeepVigor

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21808 2025-10-28 cs.CV cs.AI 57%

Semantic Relation-Enhanced CLIP Adapter for Domain Adaptive Zero-Shot Learning

Jiaao Yu, Mingjie Han, Jinkun Jiang, Junyu Dong, Tao Gong, Man Lan

机构 * School of Computer Science and Technology, East China Normal University, China(东华大学计算机科学与技术学院) College of Computer Science and Technology, Ocean University of China, China(中国海洋大学计算机科学与技术学院)

专题命中 其他安全 :alignment(abstract);分类 cs.AI

Comments 5 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21806 2025-10-28 cs.CV cs.AI 57%

Frame-Difference Guided Dynamic Region Perception for CLIP Adaptation in Text-Video Retrieval

Jiaao Yu, Mingjie Han, Tao Gong, Jian Zhang, Man Lan

机构 * School of Computer Science and Technology, East China Normal University, China(上海师范大学计算机科学与技术学院) School of Information Science and Technology, University of Science and Technology of China(中国科学技术大学信息科学与技术学院)

专题命中 其他安全 :alignment(abstract);分类 cs.AI

Comments 5 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22754 2025-10-28 cs.RO 50%

TWC-SLAM: Multi-Agent Cooperative SLAM with Text Semantics and WiFi Features Integration for Similar Indoor Environments

Chunyu Li, Shoubin Chen, Dong Li, Weixing Xue, Qingquan Li

机构 * Guangdong Laboratory of Artificial Intelligence and Digital Economy (SZ)(广东人工智能与数字经济实验室) Shenzhen University(深圳大学) College of natural resources and environment(自然资源与环境学院)

专题命中 其他安全 :alignment(abstract)

Comments Accepted by the IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.19333 2025-10-28 cs.CV 50%

A Training-Free Framework for Open-Vocabulary Image Segmentation and Recognition with EfficientNet and CLIP

Ying Dai, Wei Yu Chen

专题命中 其他安全 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.00289 2025-10-28 physics.bio-ph q-bio.PE 50%

Phylogenetic Corrections and Higher-Order Sequence Statistics in Protein Families: The Potts Model vs MSA Transformer

Kisan Khatri, Ronald M. Levy, Allan Haldane

专题命中 其他安全 :alignment(abstract)

Comments 7 pages, 5 figures, Also presented in BPS2025 Annual Meeting, Los Angeles, California

详情

展开后加载摘要…

URL PDF HTML 收藏