arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 9400 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 9400 篇

2508.15526 2025-08-22 cs.CL 79%

SafetyFlow: An Agent-Flow System for Automated LLM Safety Benchmarking

Xiangyang Zhu, Yuan Tian, Chunyi Li, Kaiwei Zhang, Wei Sun, Guangtao Zhai

机构 * Shanghai AI Lab(上海人工智能实验室) East China Normal University(东华大学)

专题命中 安全评测 :safety(title,abstract);分类 cs.CL

Comments Code and dataset are available at https://github.com/yangyangyang127/SafetyFlow

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.14741 2025-08-21 cs.LG 79%

CaTE Data Curation for Trustworthy AI

Mary Versa Clemens-Sewall, Christopher Cervantes, Emma Rafkin, J. Neil Otte, Tom Magelinski, Libby Lewis, Michelle Liu, Dana Udwin, Monique Kirkman-Bey

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.13787 2025-08-20 cs.MA cs.AI cs.NI 79%

BetaWeb: Towards a Blockchain-enabled Trustworthy Agentic Web

Zihan Guo, Yuanjian Zhou, Chenyi Wang, Linlin You, Minjie Bian, Weinan Zhang

机构 * Shanghai Innovation Institute(上海创新研究院) Sun Yat-sen University(中山大学) Zhejiang University(浙江大学) Shanghai Data Group Co., Ltd(上海数据集团有限公司) Shanghai Jiao Tong University(上海交通大学)

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI

Comments A technical report with 21 pages, 3 figures, and 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.10872 2025-08-19 cs.CV cs.AI cs.ET 79%

V-RoAst: Visual Road Assessment. Can VLM be a Road Safety Assessor Using the iRAP Standard?

Natchapon Jongwiriyanurak, Zichao Zeng, June Moh Goo, Xinglei Wang, Ilya Ilyankou, Kerkritt Sriroongvikrai, Nicola Christie, Meihui Wang, Huanfa Chen, James Haworth

机构 * University College London(伦敦大学学院) Chulalongkorn University(朱拉隆梭大学)

专题命中 安全评测 :safety(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11406 2025-08-18 cs.RO cs.AI 79%

Open, Reproducible and Trustworthy Robot-Based Experiments with Virtual Labs and Digital-Twin-Based Execution Tracing

Benjamin Alt, Mareike Picklum, Sorin Arion, Franklin Kenghagho Kenfack, Michael Beetz

机构 * AICOR Institute for Artificial Intelligence, University of Bremen(人工智能研究所,不莱梅大学)

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI

Comments 8 pages, 6 figures, submitted to the 1st IROS Workshop on Embodied AI and Robotics for Future Scientific Discovery

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05148 2025-08-14 cs.RO cs.AI 79%

Chemist Eye: A Visual Language Model-Powered System for Safety Monitoring and Robot Decision-Making in Self-Driving Laboratories

Francisco Munguia-Galeano, Zhengxue Zhou, Satheeshkumar Veeramani, Hatem Fakhruldeen, Louis Longley, Rob Clowes, Andrew I. Cooper

机构 * Cooper Group, Department of Chemistry, University of Liverpool(Cooper集团,化学系,利物浦大学)

专题命中 安全评测 :safety(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.23486 2025-08-14 cs.CL 79%

A Novel Evaluation Benchmark for Medical LLMs: Illuminating Safety and Effectiveness in Clinical Domains

Shirui Wang, Zhihui Tang, Huaxia Yang, Qiuhong Gong, Tiantian Gu, Hongyang Ma, Yongxin Wang, Wubin Sun, Zeliang Lian, Kehang Mao, Yinan Jiang, Zhicheng Huang, Lingyun Ma, Wenjie Shen, Yajie Ji, Yunhui Tan, Chunbo Wang, Yunlu Gao, Qianling Ye, Rui Lin, Mingyu Chen, Lijuan Niu, Zhihao Wang, Peng Yu, Mengran Lang, Yue Liu, Huimin Zhang, Haitao Shen, Long Chen, Qiguang Zhao, Si-Xuan Liu, Lina Zhou, Hua Gao, Dongqiang Ye, Lingmin Meng, Youtao Yu, Naixin Liang, Jianxiong Wu

专题命中 安全评测 :safety(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.07866 2025-08-14 cs.AI 79%

System 2 Reasoning for Human-AI Alignment: Generality and Adaptivity via ARC-AGI

Sejin Kim, Sundong Kim

专题命中 安全评测 :alignment(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07390 2025-08-12 cs.HC cs.AI 79%

Urbanite: A Dataflow-Based Framework for Human-AI Interactive Alignment in Urban Visual Analytics

Gustavo Moreira, Leonardo Ferreira, Carolina Veiga, Maryam Hosseini, Fabio Miranda

机构 * University of Illinois Chicago(伊利诺伊大学芝加哥分校) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) University of California, Berkeley(加州大学伯克利分校)

专题命中 安全评测 :alignment(title,abstract);分类 cs.AI

Comments Accepted at IEEE VIS 2025. Urbanite is available at https://urbantk.org/urbanite

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07031 2025-08-12 eess.IV cs.AI cs.CV 79%

Trustworthy Medical Imaging with Large Language Models: A Study of Hallucinations Across Modalities

Anindya Bijoy Das, Shahnewaz Karim Sakib, Shibbir Ahmed

机构 * The University of Akron(阿克隆大学) University of Tennessee at Chattanooga(田纳西大学查塔努加分校) Texas State University(德克萨斯州立大学)

专题命中 安全评测 :trustworthy(title);safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06885 2025-08-12 cs.LG 79%

Conformal Prediction and Trustworthy AI

Anthony Bellotti, Xindi Zhao

机构 * School of Computer Science, University of Nottingham Ningbo China(诺丁汉大学宁波校区计算机科学学院)

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.LG

Comments Preprint for an essay to be published in The Importance of Being Learnable (Enhancing the Learnability and Reliability of Machine Learning Algorithms) Essays Dedicated to Alexander Gammerman on His 80th Birthday, LNCS Springer Nature Switzerland AG ed. Nguyen K.A. and Luo Z

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05855 2025-08-11 cs.AI cs.RO 79%

Safety of Embodied Navigation: A Survey

Zixia Wang, Jia Hu, Ronghui Mu

机构 * University of Exeter(埃克塞特大学)

专题命中 安全评测 :safety(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04251 2025-08-07 cs.LG 79%

T3Time: Tri-Modal Time Series Forecasting via Adaptive Multi-Head Alignment and Residual Fusion

Abdul Monaf Chowdhury, Rabeya Akter, Safaeid Hossain Arib

专题命中 安全评测 :alignment(title,abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03117 2025-08-06 cs.AI 79%

Toward a Trustworthy Optimization Modeling Agent via Verifiable Synthetic Data Generation

Vinicius Lima, Dzung T. Phan, Jayant Kalagnanam, Dhaval Patel, Nianjun Zhou

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI

Comments 25 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21782 2025-07-30 cs.CL 79%

The Problem with Safety Classification is not just the Models

Sowmya Vajjala

专题命中 安全评测 :safety(title,abstract);分类 cs.CL

Comments Pre-print, Short paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14900 2025-07-24 cs.CL 79%

From Neurons to Semantics: Evaluating Cross-Linguistic Alignment Capabilities of Large Language Models via Neurons Alignment

Chongxuan Huang, Yongshi Ye, Biao Fu, Qifeng Su, Xiaodong Shi

机构 * School of Informatics, Xiamen University(厦门大学信息学院) Institute of Artificial Intelligence, Xiamen University(厦门大学人工智能学院) Key Laboratory of Digital Protection and Intelligent Processing of Intangible Cultural Heritage of Fujian and Taiwan (Xiamen University), Ministry of Culture and Tourism(福建省和台湾非物质文化遗产数字化保护与智能处理重点实验室(厦门大学),文化和旅游部)

专题命中 安全评测 :alignment(title,abstract);分类 cs.CL

Comments ACL main 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.16033 2025-07-23 cs.HC cs.AI 79%

"Just a strange pic": Evaluating 'safety' in GenAI Image safety annotation tasks from diverse annotators' perspectives

Ding Wang, Mark Díaz, Charvi Rastogi, Aida Davani, Vinodkumar Prabhakaran, Pushkar Mishra, Roma Patel, Alicia Parrish, Zoe Ashwood, Michela Paganini, Tian Huey Teh, Verena Rieser, Lora Aroyo

专题命中 安全评测 :safety(title,abstract);分类 cs.AI

Comments Accepted to AAAI/ACM Conference on Artificial Intelligence, Ethics, and Society 2025 (AIES 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14693 2025-07-22 cs.CL cs.AI cs.CY cs.LG 79%

Rethinking Suicidal Ideation Detection: A Trustworthy Annotation Framework and Cross-Lingual Model Evaluation

Amina Dzafic, Merve Kavut, Ulya Bayram

专题命中 安全评测 :trustworthy(title);分类 cs.CL、cs.AI、cs.CY

Comments This manuscript has been submitted to the IEEE Journal of Biomedical and Health Informatics

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.19819 2025-07-17 cs.AR cs.AI 79%

ChipAlign: Instruction Alignment in Large Language Models for Chip Design via Geodesic Interpolation

Chenhui Deng, Yunsheng Bai, Haoxing Ren

机构 * NVIDIA(英伟达)

专题命中 安全评测 :alignment(title,abstract);分类 cs.AI

Comments Accepted to DAC 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.10905 2025-07-15 cs.HC cs.AI cs.ET 79%

The impact of labeling automotive AI as "trustworthy" or "reliable" on user evaluation and technology acceptance

John Dorsch, Ophelia Deroy

机构 * Faculty of Philosophy, Philosophy of Science and the Study of Religion, Ludwig-Maximilians-Universität München(哲学学院、科学哲学与宗教研究学院,慕尼黑路德维希-马克西米利安大学)

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI

Comments 36 pages, 12 figures

Journal ref Nature Scientific Reports, 15, 1481. 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.08109 2025-07-14 cs.CL 79%

Audit, Alignment, and Optimization of LM-Powered Subroutines with Application to Public Comment Processing

Reilly Raab, Mike Parker, Dan Nally, Sadie Montgomery, Anastasia Bernat, Sai Munikoti, Sameera Horawalavithana

机构 * Pacific Northwest National Laboratory(太平洋西北国家实验室)

专题命中 安全评测 :alignment(title);safety(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.07787 2025-07-14 cs.AI 79%

Measuring AI Alignment with Human Flourishing

Elizabeth Hilliard, Akshaya Jagadeesh, Alex Cook, Steele Billings, Nicholas Skytland, Alicia Llewellyn, Jackson Paull, Nathan Paull, Nolan Kurylo, Keatra Nesbitt, Robert Gruenewald, Anthony Jantzi, Omar Chavez

机构 * Gloo Valkyrie Intelligence Biblica

专题命中 安全评测 :alignment(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.06378 2025-07-10 cs.CL 79%

Evaluating Morphological Alignment of Tokenizers in 70 Languages

Catherine Arnett, Marisa Hudspeth, Brendan O'Connor

机构 * University of Massachusetts Amherst(马萨诸塞大学阿默斯特分校)

专题命中 安全评测 :alignment(title,abstract);分类 cs.CL

Comments 6 pages, 3 figures. Accepted to the Tokenization Workshop at ICML 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.06949 2025-07-09 cs.SE cs.CL 79%

Towards Exception Safety Code Generation with Intermediate Representation Agents Framework

Xuanming Zhang, Yuxuan Chen, Yuan Yuan, Minlie Huang

机构 * University of Wisconsin-Madison(威斯康星大学麦迪逊分校) Stanford University(斯坦福大学) Tsinghua University(清华大学) Beihang University(北航大学)

专题命中 安全评测 :safety(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.04976 2025-07-08 cs.CV cs.CL 79%

Can Video LLMs Refuse to Answer? Alignment for Answerability in Video Large Language Models

Eunseop Yoon, Hee Suk Yoon, Mark A. Hasegawa-Johnson, Chang D. Yoo

机构 * Korea Advanced Institute of Science and Technology (KAIST)(韩国科学技术院) University of Illinois at Urbana-Champaign (UIUC)(伊利诺伊大学厄巴纳-香槟分校)

专题命中 安全评测 :alignment(title,abstract);分类 cs.CL

Comments ICLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.01274 2025-07-03 cs.HC cs.AI 79%

AI Meets Maritime Training: Precision Analytics for Enhanced Safety and Performance

Vishakha Lall, Yisi Liu

专题命中 安全评测 :safety(title,abstract);分类 cs.AI

Comments Accepted and Presented at 11th International Maritime Science Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.00096 2025-07-02 cs.CR cs.AI 79%

AI-Governed Agent Architecture for Web-Trustworthy Tokenization of Alternative Assets

Ailiya Borjigin, Wei Zhou, Cong He

机构 * Probe Group Pte. Ltd.(Probe集团)

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI

Comments 8 Pages, 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.20832 2025-07-01 cs.CV cs.AI 79%

Leveraging Vision-Language Models to Select Trustworthy Super-Resolution Samples Generated by Diffusion Models

Cansu Korkmaz, Ahmet Murat Tekalp, Zafer Dogan

机构 * Department of Electrical and Electronics Engineering and KUIS AI Center, Koç University(电气与电子工程系和KUIS人工智能中心,科克大学)

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI

Comments 14 pages, 9 figures, 5 tables, accepted to IEEE Transactions on Circuits and Systems for Video Technology

Journal ref IEEE Transactions on Circuits and Systems for Video Technology 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.17404 2025-07-01 cs.AI 79%

Super Co-alignment of Human and AI for Sustainable Symbiotic Society

Yi Zeng, Feifei Zhao, Yuwei Wang, Enmeng Lu, Yaodong Yang, Lei Wang, Chao Liu, Yitao Liang, Dongcheng Zhao, Bing Han, Haibo Tong, Yao Liang, Dongqi Liang, Kang Sun, Boyuan Chen, Jinyu Fan

机构 * Beijing Key Laboratory of Safe AI and Superalignment(北京安全人工智能与超对齐关键实验室) Beijing Institute of AI Safety and Governance(北京人工智能安全与治理研究院) Brain-inspired Cognitive AI Lab(脑启发认知人工智能实验室) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) University of Chinese Academy of Sciences(中国科学院大学) Wenge Technology Co., Ltd.(Wenger 技术有限公司) Institute for Artificial Intelligence, Peking University(北京大学人工智能研究院) Long-term AI(长期人工智能) State Key Laboratory of Cognitive Neuroscience and Learning, Beijing Normal University(北京师范大学认知神经科学与学习国家重点实验室)

专题命中 安全评测 :alignment(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.13232 2025-06-30 cs.AI cs.CV 79%

StarFT: Robust Fine-tuning of Zero-shot Models via Spuriosity Alignment

Younghyun Kim, Jongheon Jeong, Sangkyung Kwak, Kyungmin Lee, Juho Lee, Jinwoo Shin

机构 * Samsung(三星) Korea University(韩国大学) General Robotics(通用机器人) KAIST(韩国科学技术院)

专题命中 安全评测 :alignment(title,abstract);分类 cs.AI

Comments IJCAI 2025; Code is available at https://github.com/alinlab/StarFT

详情

展开后加载摘要…

URL PDF HTML 收藏