arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 9346 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 9346 篇

2504.21034 2025-09-01 cs.CR cs.AI cs.LG 62%

SAGA: A Security Architecture for Governing AI Agentic Systems

Georgios Syros, Anshuman Suri, Jacob Ginesin, Cristina Nita-Rotaru, Alina Oprea

机构 * Northeastern University(东北大学)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.21389 2025-09-01 cs.CL cs.AI 62%

AllSummedUp: un framework open-source pour comparer les metriques d'evaluation de resume

Tanguy Herserant, Vincent Guigue

机构 * AgroParisTech - MIA(阿格罗巴黎技术学院-信息分析中心)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

Comments in French language

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.21061 2025-08-29 cs.HC cs.AI cs.LG 62%

OnGoal: Tracking and Visualizing Conversational Goals in Multi-Turn Dialogue with Large Language Models

Adam Coscia, Shunan Guo, Eunyee Koh, Alex Endert

机构 * Georgia Institute of Technology(佐治亚理工学院) Adobe Research(Adobe研究)

专题命中 安全评测 :alignment(abstract);分类 cs.AI、cs.LG

Comments Accepted to UIST 2025. 18 pages, 9 figures, 2 tables. For a demo video, see https://youtu.be/uobhmxo6EIE

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20416 2025-08-29 cs.CL cs.AI 62%

DentalBench: Benchmarking and Advancing LLMs Capability for Bilingual Dentistry Understanding

Hengchuan Zhu, Yihuan Xu, Yichen Li, Zijie Meng, Zuozhu Liu

机构 * Zhejiang University(浙江大学) ZJU-Angelalign R&D Center for Intelligence Healthcare(浙江大学智能医疗研发中心)

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.19487 2025-08-28 cs.LG cs.AI 62%

Data-Efficient Symbolic Regression via Foundation Model Distillation

Wangyang Ying, Jinghan Zhang, Haoyue Bai, Nanxu Gong, Xinyuan Wang, Kunpeng Liu, Chandan K. Reddy, Yanjie Fu

机构 * Institute for Clarity in Documentation(文档清晰研究所) Inria Paris-Rocquencourt(巴黎-罗克琴特研究所) Rajiv Gandhi University(拉吉夫·甘地大学) Tsinghua University(清华大学) Palmer Research Laboratories(帕尔默研究实验室) Arizona State University(亚利桑那州立大学) Clemson University(克莱姆森大学)

专题命中 安全评测 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.19271 2025-08-28 cs.CL cs.AI 62%

Rethinking Reasoning in LLMs: Neuro-Symbolic Local RetoMaton Beyond ICL and CoT

Rushitha Santhoshi Mamidala, Anshuman Chhabra, Ankur Mali

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.18891 2025-08-27 cs.LG cs.AI 62%

pyFAST: A Modular PyTorch Framework for Time Series Modeling with Multi-source and Sparse Data

Zhijin Wang, Senzhen Wu, Yue Hu, Xiufeng Liu

机构 * College of Computer Engineering, Jimei University(嘉应大学计算机工程学院) Chengyi College, Jimei University(嘉应大学 Chengyi 学院) Department of Technology, Management and Economics, Technical University of Denmark(丹麦技术大学技术、管理与经济系)

专题命中 安全评测 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.08846 2025-08-27 cs.CR cs.CL cs.LG 62%

Fingerprint Vector: Enabling Scalable and Efficient Model Fingerprint Transfer via Vector Addition

Zhenhua Xu, Qichen Liu, Zhebo Wang, Wenpeng Xing, Dezhang Kong, Mohan Li, Meng Han

专题命中 安全评测 :harmlessness(abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17786 2025-08-26 cs.AI cs.FL cs.LG cs.LO 62%

Interpretable Early Failure Detection via Machine Learning and Trace Checking-based Monitoring

Andrea Brunello, Luca Geatti, Angelo Montanari, Nicola Saccomanno

机构 * University of Udine, Italy(乌迪内大学)

专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.LG

Comments Full version of the paper accepted for publication at the 28th European Conference on Artificial Intelligence (ECAI 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17680 2025-08-26 cs.LG cs.AI cs.CV 62%

Robustness Feature Adapter for Efficient Adversarial Training

Quanwei Wu, Jun Guo, Wei Wang, Yi Wang

机构 * Dongguan University of Technology(东莞科技学院) The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州))

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG

Comments The paper has been accepted for presentation at ECAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11280 2025-08-26 cs.CL cs.AI 62%

LETToT: Label-Free Evaluation of Large Language Models On Tourism Using Expert Tree-of-Thought

Ruiyan Qi, Congding Wen, Weibo Zhou, Jiwei Li, Shangsong Liang, Lingbo Li

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.14745 2025-08-26 cs.LG cs.AI 62%

Explainable Prediction of the Mechanical Properties of Composites with CNNs

Varun Raaghav, Dimitrios Bikos, Antonio Rago, Francesca Toni, Maria Charalambides

机构 * Department of Mechanical Engineering, Imperial College London, UK(帝国理工学院机械工程系) Department of Computing, Imperial College London, UK(帝国理工学院计算系) Department of Informatics, King's College London, UK(伦敦国王学院信息学系)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG

Comments 9 pages, 6 figures. Accepted for publication at The 14th Conference on Prestigious Applications of Intelligent Systems (PAIS-2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.16672 2025-08-26 cs.CY cs.AI 62%

The AI Model Risk Catalog: What Developers and Researchers Miss About Real-World AI Harms

Pooja S. B. Rao, Sanja Šćepanović, Dinesh Babu Jayagopi, Mauro Cherubini, Daniele Quercia

专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.CY

Comments Accepted to AIES 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.16669 2025-08-26 cs.CY cs.AI cs.HC 62%

Situational Awareness as the Imperative Capability for Disaster Resilience in the Era of Complex Hazards and Artificial Intelligence

Hongrak Pak, Ali Mostafavi

机构 * organization= UrbanResilience.AI Lab, Zachry Department of Civil Environmental Engineering, Texas A\&M University , city= College Station , postcode= 77843 , state= TX , country= USA

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15910 2025-08-25 cs.CL cs.AI cs.IR 62%

Evaluating Structured Decoding for Text-to-Table Generation: Evidence from Three Datasets

Julian Oestreich, Lydia Müller

机构 * Institute for Applied Informatics (InfAI) at Leipzig University(应用信息学院(InfAI))

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

Comments to be published in the workshop proceedings of the "From Rules to Language Models: Comparative Performance Evaluation" workshop, held alongside RANLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15876 2025-08-25 cs.CL cs.AI cs.MA 62%

DeepMEL: A Multi-Agent Collaboration Framework for Multimodal Entity Linking

Fang Wang, Tianwei Yan, Zonghao Yang, Minghao Hu, Jun Zhang, Zhunchen Luo, Xiaoying Bai

机构 * School of Computer Science, Peking University(北京大学计算机科学学院) School of Information Science and Engineering, Chongqing Jiaotong University(重庆交通大学信息科学与工程学院) China Research and Development Academy of Machinery Equipment(机械电子研究发展院) Center of Information Research, Academy of Military Science(军事科学信息研究中心)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15853 2025-08-25 cs.CL cs.AI cs.SD eess.AS 62%

MGSC: A Multi-granularity Consistency Framework for Robust End-to-end Asr

Xuwen Yang

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

Comments 12 pages, 5figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15464 2025-08-22 cs.CL cs.AI 62%

RadReason: Radiology Report Evaluation Metric with Reasons and Sub-Scores

Yingshu Li, Yunyi Liu, Lingqiao Liu, Lei Wang, Luping Zhou

机构 * School of Electrical and Computer Engineering, University of Sydney(悉尼大学电气与计算机工程学院) School of Computer Science, University of Adelaide(阿德莱德大学计算机科学学院) School of Computing and Information Technology, University of Wollongong(沃林根大学计算与信息科技学院)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15220 2025-08-22 cs.LG cs.AI cs.LO 62%

Locally Pareto-Optimal Interpretations for Black-Box Machine Learning Models

Aniruddha Joshi, Supratik Chakraborty, S Akshay, Shetal Shah, Hazem Torfah, Sanjit Seshia

机构 * University of California at Berkeley(加州大学伯克利分校) Indian Institute of Technology Bombay(印度班加罗尔理工学院) Chalmers University of Technology and University of Gothenburg(查尔姆斯理工大学和哥德堡大学)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG

Comments This work has been accepted at ATVA'25

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.14921 2025-08-22 cs.CY cs.AI 62%

Designing an Interdisciplinary Artificial Intelligence Curriculum for Engineering: Evaluation and Insights from Experts

Johannes Schleiss, Anke Manukjan, Michelle Ines Bieber, Sebastian Lang, Sebastian Stober

机构 * Otto von Guericke University Magdeburg(奥托·冯·格里克大学马格德堡)

专题命中 安全评测 :alignment(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.04742 2025-08-22 cs.CY cs.AI 62%

A Case for Specialisation in Non-Human Entities

El-Mahdi El-Mhamdi, Lê-Nguyên Hoang, Mariame Tighanimine

专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.CY

Comments Accepted to AAAI/ACM AIES 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.19422 2025-08-21 cs.CY cs.AI cs.HC 62%

Generative AI in K-12 Education: The CyberScholar Initiative

Vania Castro, Ana Karina de Oliveira Nascimento, Raigul Zheldibayeva, Duane Searsmith, Akash Saini, Bill Cope, Mary Kalantzis

专题命中 安全评测 :alignment(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.13180 2025-08-20 cs.AI cs.LG 62%

Search-Time Data Contamination

Ziwen Han, Meher Mankikar, Julian Michael, Zifan Wang

机构 * Scale AI

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.13152 2025-08-19 cs.CL cs.AI 62%

RepreGuard: Detecting LLM-Generated Text by Revealing Hidden Representation Patterns

Xin Chen, Junchao Wu, Shu Yang, Runzhe Zhan, Zeyu Wu, Ziyang Luo, Di Wang, Min Yang, Lidia S. Chao, Derek F. Wong

机构 * NLP(自然语言处理) CT Lab, Department of Computer and Information Science, University of Macau(计算机与信息科学系,澳门大学) Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences(深圳先进技术研究院,中国科学院) Provable Responsible AI and Data Analytics Lab, KAUST(可证明责任AI与数据分析实验室,卡尔斯兰大学) Hong Kong Baptist University(香港 Baptist 大学)

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL、cs.AI

Comments Accepted to TACL 2025. This version is a pre-MIT Press publication version

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12212 2025-08-19 cs.LG cs.AI q-bio.QM 62%

ProtTeX-CC: Activating In-Context Learning in Protein LLM via Two-Stage Instruction Compression

Chuanliu Fan, Zicheng Ma, Jun Gao, Nan Yu, Jun Zhang, Ziqiang Cao, Yi Qin Gao, Guohong Fu

专题命中 安全评测 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.04596 2025-08-19 cs.AI cs.CE cs.CL 62%

SECQUE: A Benchmark for Evaluating Real-World Financial Analysis Capabilities

Noga Ben Yoash, Meni Brief, Oded Ovadia, Gil Shenderovitz, Moshik Mishaeli, Rachel Lemberg, Eitam Sheetrit

机构 * Microsoft Industry AI(微软产业人工智能)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

Comments Benchmark available at: https://huggingface.co/datasets/nogabenyoash/SecQue

Journal ref n Proceedings of the Fourth Workshop on Generation, Evaluation and Metrics, Association for Computational Linguistics (2025) https://aclanthology.org/2025.gem-1.16/

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10494 2025-08-15 cs.LG cs.AI cs.MA 62%

A Unified Multi-Agent Framework for Universal Multimodal Understanding and Generation

Jiulin Li, Ping Huang, Yexin Li, Shuo Chen, Juewen Hu, Ye Tian

专题命中 安全评测 :alignment(abstract);分类 cs.AI、cs.LG

Comments 8 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10433 2025-08-15 cs.AI cs.CV cs.LG 62%

We-Math 2.0: A Versatile MathBook System for Incentivizing Visual Mathematical Reasoning

Runqi Qiao, Qiuna Tan, Peiqing Yang, Yanzi Wang, Xiaowan Wang, Enhui Wan, Sitong Zhou, Guanting Dong, Yuchen Zeng, Yida Xu, Jie Wang, Chong Sun, Chen Li, Honggang Zhang

机构 * BUPT(北京邮电大学) WeChat Vision, Tencent Inc.(腾讯公司) Tsinghua University(清华大学)

专题命中 安全评测 :alignment(abstract);分类 cs.AI、cs.LG

Comments Working in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10022 2025-08-15 cs.CL cs.AI 62%

Conformal P-Value in Multiple-Choice Question Answering Tasks with Provable Risk Control

Yuanchang Ye

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.13871 2025-08-15 cs.LG cs.AI cs.CR 62%

An Explainable Transformer-based Model for Phishing Email Detection: A Large Language Model Approach

Mohammad Amaz Uddin, Md Mahiuddin, Iqbal H. Sarker

机构 * Department of Computer Science and Engineering, BGC Trust University Bangladesh(Bangladesh BGC Trust 大学 计算机科学与工程系) Department of Computer Science and Engineering, International Islamic University Chittagong(Bangladesh 国际伊斯兰大学 昌德加荣分校 计算机科学与工程系) Centre for Securing Digital Futures, School of Science, Edith Cowan University(澳大利亚 埃德温·考文大学 科学学院 安全数字未来中心)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏