arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-11-11 至 2025-11-11 共收录 30 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 30 篇

2511.06890 2025-11-11 cs.CL 87%

EduGuardBench: A Holistic Benchmark for Evaluating the Pedagogical Fidelity and Adversarial Safety of LLMs as Simulated Teachers

Yilin Jiang, Mingzi Zhang, Xuanyu Yin, Sheng Jin, Suyu Lu, Zuocan Ying, Zengyi Yu, Xiangjie Kong

专题命中 安全评测 :safety(title,abstract);alignment(abstract);trustworthy(abstract);AI safety(abstract)

Comments 22 pages, 9 figures, accepted by AAAI2026 as oral paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.05889 2025-11-11 cs.RO 82%

From Words to Safety: Language-Conditioned Safety Filtering for Robot Navigation

Zeyuan Feng, Haimingyue Zhang, Somil Bansal

机构 * department of Aeronautics and Astronautics, Stanford University(航空与宇航系,斯坦福大学) School of Vehicle and Mobility at Tsinghua University(清华大学车辆与移动系统学院)

专题命中 安全评测 :safety(title,abstract);trustworthy(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01710 2025-11-11 cs.CL cs.LG 81%

CultureGuard: Towards Culturally-Aware Dataset and Guard Model for Multilingual Safety Applications

Raviraj Joshi, Rakesh Paul, Kanishk Singla, Anusha Kamath, Michael Evans, Katherine Luna, Shaona Ghosh, Utkarsh Vaidya, Eileen Long, Sanjay Singh Chauhan, Niranjan Wartikar

机构 * NVIDIA(英伟达)

专题命中 安全评测 :safety(title,abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.11881 2025-11-11 cs.CL 79%

Evaluating Human-LLM Representation Alignment: A Case Study on Affective Sentence Generation for Augmentative and Alternative Communication

Shadab Choudhury, Asha Kumar, Lara J. Martin

专题命中 安全评测 :alignment(title,abstract);分类 cs.CL

Comments Published at IJCNLP-AACL 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.05526 2025-11-11 cs.CY 77%

Emergency Response Measures for Catastrophic AI Risk

James Zhang, Miles Kodama, Zongze Wu, Michael Chen, Yue Zhu, Geng Hong

专题命中 安全评测 :alignment(abstract);safety(abstract);AI safety(abstract);分类 cs.CY

Comments Accepted to the Workshop on Regulatable ML at the 39th Conference on Neural Information Processing Systems (NeurIPS 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.05682 2025-11-11 cs.CV cs.LG 70%

VMDT: Decoding the Trustworthiness of Video Foundation Models

Yujin Potter, Zhun Wang, Nicholas Crispino, Kyle Montgomery, Alexander Xiong, Ethan Y. Chang, Francesco Pinto, Yuqi Chen, Rahul Gupta, Morteza Ziyadi, Christos Christodoulopoulos, Bo Li, Chenguang Wang, Dawn Song

机构 * University of California, Berkeley(加州大学伯克利分校) University of California, Santa Cruz(加州大学圣克ruz分校) University of Illinois at Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) University of Chicago(芝加哥大学) Amazon(亚马逊) Information Commissioner’s Office(信息专员办公室)

专题命中 安全评测 :safety(abstract);trustworthy(abstract);分类 cs.LG

Comments NeurIPS 2025 Datasets & Benchmarks

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.13333 2025-11-11 cs.AI 70%

Evaluation Awareness Scales Predictably in Open-Weights Large Language Models

Maheep Chaudhary, Ian Su, Nikhil Hooda, Nishith Shankar, Julia Tan, Kevin Zhu, Ryan Lagasse, Vasu Sharma, Ashwinee Panda

机构 * Independent(独立研究者) UC Irvine(加州大学尔湾分校) University of Waterloo(滑铁卢大学) UW Madison(威斯康星大学麦迪逊分校) University of Toronto(多伦多大学) Algoverse(Algoverse公司) META University of Maryland(马里兰大学)

专题命中 安全评测 :safety(abstract);AI safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.07070 2025-11-11 cs.AI cs.LG 62%

RedOne 2.0: Rethinking Domain-specific LLM Post-Training in Social Networking Services

Fei Zhao, Chonggang Lu, Haofu Qian, Fangcheng Shi, Zijie Meng, Jianzhao Huang, Xu Tang, Zheyong Xie, Zheyu Ye, Zhe Xu, Yao Hu, Shaosheng Cao

机构 * NLP Team, Xiaohongshu Inc.(小红书研究院自然语言处理团队)

专题命中 安全评测 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.06763 2025-11-11 cs.CL cs.AI 62%

Sensitivity of Small Language Models to Fine-tuning Data Contamination

Nicy Scaria, Silvester John Joseph Kennedy, Deepak Subramani

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.05527 2025-11-11 cs.CL cs.AI 62%

Bridging Industrial Expertise and XR with LLM-Powered Conversational Agents

Despina Tomkou, George Fatouros, Andreas Andreou, Georgios Makridis, Fotis Liarokapis, Dimitrios Dardanis, Athanasios Kiourtis, John Soldatos, Dimosthenis Kyriazis

机构 * Innov-Acts Ltd.(Innov-Acts有限公司) CYENS Centre of Excellence(CYENS卓越中心) University of Piraeus(比雷埃克斯大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

Comments 7 pages, 7 figures

Journal ref 2025 21st International Conference on Distributed Computing in Smart Systems and the Internet of Things (DCOSS-IoT)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.06470 2025-11-11 cs.AI cs.LG 62%

Brain-Inspired Planning for Better Generalization in Reinforcement Learning

Mingde "Harry" Zhao

专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.LG

Comments McGill PhD Thesis (updated on 20251109 for typos and margin adjustments)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.06248 2025-11-11 cs.LG cs.AI 62%

Constraint-Informed Active Learning for End-to-End ACOPF Optimization Proxies

Miao Li, Michael Klamkin, Pascal Van Hentenryck, Wenting Li, Russell Bent

机构 * University of Texas at Austin, TX, USA(德克萨斯大学奥斯汀分校)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG

Comments 8 PAGES

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.05516 2025-11-11 cs.CL cs.AI cs.SD eess.AS 62%

Ming-UniAudio: Speech LLM for Joint Understanding, Generation and Editing with Unified Representation

Canxiang Yan, Chunxiang Jin, Dawei Huang, Haibing Yu, Han Peng, Hui Zhan, Jie Gao, Jing Peng, Jingdong Chen, Jun Zhou, Kaimeng Ren, Ming Yang, Mingxue Yang, Qiang Xu, Qin Zhao, Ruijie Xiong, Shaoxiong Lin, Xuezhi Wang, Yi Yuan, Yifei Wu, Yongjie Lyu, Zhengyu He, Zhihao Qiu, Zhiqiang Fang, Ziyuan Huang

机构 * Inclusion AI Ant Group(Inclusion AI Ant集团)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

Comments 32 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00710 2025-11-11 cs.AI cs.CL 62%

On Verifiable Legal Reasoning: A Multi-Agent Framework with Formalized Knowledge Representations

Albert Sadowski, Jarosław A. Chudziak

机构 * Warsaw University of Technology(华沙技术大学)

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL、cs.AI

Comments Accepted for publication at the 34th ACM International Conference on Information and Knowledge Management (CIKM '25)

Journal ref CIKM '25: Proceedings of the 34th ACM International Conference on Information and Knowledge Management (2025) 2535-2545

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.07298 2025-11-11 cs.CV cs.AI 57%

LMM-IQA: Image Quality Assessment for Low-Dose CT Imaging

Kagan Celik, Mehmet Ozan Unal, Metin Ertas, Isa Yildirim

机构 * Department of Electronics and Communication Engineering, Istanbul Technical University(电子与通信工程系,伊斯坦布尔技术大学)

专题命中 安全评测 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.07032 2025-11-11 cs.LG stat.ML 57%

Fair Bayesian Data Selection via Generalized Discrepancy Measures

Yixuan Zhang, Jiabin Luo, Zhenggang Wang, Feng Zhou, Quyu Kong

专题命中 安全评测 :alignment(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.06530 2025-11-11 cs.CL 57%

Better Datasets Start From RefineLab: Automatic Optimization for High-Quality Dataset Refinement

Xiaonan Luo, Yue Huang, Ping He, Xiangliang Zhang

机构 * University of Notre Dame(诺特大学) Vanderbilt University(范德比大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.06227 2025-11-11 cs.SE cs.AI cs.CE 57%

Assertion-Aware Test Code Summarization with Large Language Models

Anamul Haque Mollah, Ahmed Aljohani, Hyunsook Do

机构 * University of North Texas(北卡罗来纳州立大学)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

Comments Accepted for publication at 2nd ACM International Conference on AI-powered Software (AIware 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20264 2025-11-11 cs.CL 57%

EMBRACE: Shaping Inclusive Opinion Representation by Aligning Implicit Conversations with Social Norms

Abeer Aldayel, Areej Alokaili

机构 * King Saud University, College of Computer and Information Sciences(沙特王后大学,计算机与信息科学学院)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

Comments Accepted, to appear IJCNLP-AACL 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.12618 2025-11-11 cs.CL 57%

OpenUnlearning: Accelerating LLM Unlearning via Unified Benchmarking of Methods and Metrics

Vineeth Dorna, Anmol Mekala, Wenlong Zhao, Andrew McCallum, Zachary C. Lipton, J. Zico Kolter, Pratyush Maini

机构 * University of Massachusetts Amherst(马萨诸塞大学阿默斯特分校) Carnegie Mellon University(卡内基梅隆大学) DatologyAI

专题命中 安全评测 :safety(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.05936 2025-11-11 cs.RO cs.AI 57%

10 Open Challenges Steering the Future of Vision-Language-Action Models

Soujanya Poria, Navonil Majumder, Chia-Yu Hung, Amir Ali Bagherzadeh, Chuan Li, Kenneth Kwok, Ziwei Wang, Cheston Tan, Jiajun Wu, David Hsu

专题命中 安全评测 :safety(abstract);分类 cs.AI

Comments AAAI 2026 (Senior Track)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.05752 2025-11-11 cs.CL 57%

Multi-Scale Feature Fusion and Graph Neural Network Integration for Text Classification with Large Language Models

Xiangchen Song, Yulin Huang, Jinxu Guo, Yuchen Liu, Yaxuan Luan

专题命中 安全评测 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.05592 2025-11-11 cs.LG 57%

GRAVER: Generative Graph Vocabularies for Robust Graph Foundation Models Fine-tuning

Haonan Yuan, Qingyun Sun, Junhua Shi, Xingcheng Fu, Bryan Hooi, Jianxin Li, Philip S. Yu

机构 * SKLCCSE, School of Computer Science and Engineering, Beihang University(北京航空航天大学信息与电子技术学院) Key Lab of Education Blockchain and Intelligent Technology, Guangxi Normal University(广西师范大学教育区块链与智能技术重点实验室) School of Computing, National University of Singapore(新加坡国立大学计算机学院) Department of Computer Science, University of Illinois, Chicago(伊利诺伊大学芝加哥分校计算机科学系)

专题命中 安全评测 :trustworthy(abstract);分类 cs.LG

Comments Accepted by the NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.05511 2025-11-11 eess.SY cs.AI cs.SY 57%

From Failure Modes to Reliability Awareness in Generative and Agentic AI System

Janet, Lin, Liangwei Zhang

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

Comments 24pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23089 2025-11-11 cs.LG cs.NI 57%

Demystifying Network Foundation Models

Sylee Beltiukov, Satyandra Guthula, Wenbo Guo, Walter Willinger, Arpit Gupta

机构 * UC Santa Barbara(加州大学圣芭芭拉分校) NIKSUN, Inc(NIKSUN公司)

专题命中 安全评测 :alignment(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02073 2025-11-11 cs.AI 57%

Large model retrieval enhancement framework for construction site risk identification

Jiawei Li, Chengye Yang, Yaochen Zhang, Weilin Sun, Lei Meng, Xiangxu Meng

专题命中 安全评测 :safety(abstract);分类 cs.AI

Comments in Chinese language

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.06876 2025-11-11 cs.CV 50%

Generating an Image From 1,000 Words: Enhancing Text-to-Image With Structured Captions

Eyal Gutflaish, Eliran Kachlon, Hezi Zisman, Tal Hacham, Nimrod Sarid, Alexander Visheratin, Saar Huberman, Gal Davidi, Guy Bukchin, Kfir Goldberg, Ron Mokady

机构 * BRIA AI(BRIA人工智能)

专题命中 安全评测 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.06651 2025-11-11 cs.CV 50%

NOVO: Bridging LLaVA and SAM with Visual-only Prompts for Reasoning Segmentation

Kyung-Yoon Yoon, Yeong-Jun Cho

机构 * Department of Artificial Intelligence Convergence(人工智能融合系)

专题命中 安全评测 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.05531 2025-11-11 q-bio.NC eess.IV 50%

Selection and Stability of Functional Connectivity Features for Classification of Brain Disorders

Aniruddha Saha, Soujanya Hazra, Sanjay Ghosh

专题命中 安全评测 :trustworthy(abstract)

Comments 10 pages, 5 figures, and 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.18004 2025-11-11 cs.CV 50%

SkinCaRe: A Multimodal Dermatology Dataset Annotated with Medical Caption and Chain-of-Thought Reasoning

Yuhao Shen, Liyuan Sun, Yan Xu, Wenbin Liu, Shuping Zhang, Shawn Afvari, Zhongyi Han, Jiaoyan Song, Yongzhi Ji, Tao Lu, Xiaonan He, Xin Gao, Juexiao Zhou

机构 * School of Data Science, The Chinese University of Hong Kong, Shenzhen (CUHK–Shenzhen)(数据科学学院,香港中文大学(深圳)) Computer Science Program, CEMSE Division, King Abdullah University of Science and Technology (KAUST)(计算机科学项目,科学与工程学院,国王 Abdullah 科学技术大学) Center of Excellence on Smart Health, KAUST(智能健康卓越中心,国王 Abdullah 科学技术大学) Center of Excellence for Generative AI, KAUST(生成式人工智能卓越中心,国王 Abdullah 科学技术大学) Department of Dermatology, Beijing AnZhen Hospital, Capital Medical University(皮肤科,北京安贞医院,首都医科大学) Department of Dermatology, Tianjin Institute of Integrative Dermatology, Tianjin Academy of Traditional Chinese Medicine Affiliated Hospital(皮肤科,天津整合皮肤科研究院,天津中医药大学附属医院) Department of Dermatology, Beijing Aerospace General Hospital(皮肤科,北京航天总医院) Department of Dermatology, The First Affiliated Hospital, Shantou University Medical College(皮肤科,汕头大学医学院第一附属医院) DermAssure, LLC(DermAssure 公司) School of Medicine, New York Medical College(医学院,纽约医学院) Capital Medical University(首都医科大学) Department of Dermatology, Second Hospital of Jilin University(皮肤科,吉林大学第二医院) Emergency Critical Care Center, Beijing AnZhen Hospital, Capital Medical University(急诊重症中心,北京安贞医院,首都医科大学)

专题命中 安全评测 :trustworthy(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏