arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 9346 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 9346 篇

2508.09987 2025-08-14 cs.CV cs.AI cs.CL 62%

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation

Junyan Ye, Dongzhi Jiang, Zihao Wang, Leqi Zhu, Zhenghao Hu, Zilong Huang, Jun He, Zhiyuan Yan, Jinghua Yu, Hongsheng Li, Conghui He, Weijia Li

机构 * Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Sun Yat-sen University(中山大学) CUHK MMLab(香港中文大学多模态实验室) Peking University(北京大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

Comments 19 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.12867 2025-08-14 eess.AS cs.AI cs.CL 62%

EmoVoice: LLM-based Emotional Text-To-Speech Model with Freestyle Text Prompting

Guanrou Yang, Chen Yang, Qian Chen, Ziyang Ma, Wenxi Chen, Wen Wang, Tianrui Wang, Yifan Yang, Zhikang Niu, Wenrui Liu, Fan Yu, Zhihao Du, Zhifu Gao, ShiLiang Zhang, Xie Chen

机构 * Shanghai Jiao Tong University(上海交通大学) Tongyi Speech Lab(通义语音实验室) Tianjin University(天津大学) Zhejiang University(浙江大学) Shanghai Jiao Tong University, Shanghai Innovation Institute(上海交通大学上海创新研究院)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

Comments Accepted at ACMMM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08777 2025-08-13 cs.IR cs.AI cs.LG 62%

Evaluating Podcast Recommendations with Profile-Aware LLM-as-a-Judge

Francesco Fabbri, Gustavo Penha, Edoardo D'Amico, Alice Wang, Marco De Nadai, Jackie Doremus, Paul Gigioli, Andreas Damianou, Oskar Stal, Mounia Lalmas

机构 * Spotify

专题命中 安全评测 :alignment(abstract);分类 cs.AI、cs.LG

Comments Accepted at RecSys '25

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08629 2025-08-13 cs.CY cs.AI 62%

Securing Educational LLMs: A Generalised Taxonomy of Attacks on LLMs and DREAD Risk Assessment

Farzana Zahid, Anjalika Sewwandi, Lee Brandon, Vimal Kumar, Roopak Sinha

机构 * University of Waikato(怀卡托大学) Deakin University(迪金大学)

专题命中 安全评测 :jailbreak(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08277 2025-08-13 cs.CL cs.LG 62%

Objective Metrics for Evaluating Large Language Models Using External Data Sources

Haoze Du, Richard Li, Edward Gehringer

机构 * Department of Computer Science(计算机科学系) North Carolina State University(北卡罗来纳州立大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.LG

Comments This version of the paper is lightly revised from the EDM 2025 proceedings for the sake of clarity

Journal ref EDM 2025 Palermo, Italy, July, 2025, pp. 489-495. International Educational Data Mining Society (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07668 2025-08-12 cs.LG cs.AI 62%

AIS-LLM: A Unified Framework for Maritime Trajectory Prediction, Anomaly Detection, and Collision Risk Assessment with Explainable Forecasting

Hyobin Park, Jinwook Jung, Minseok Seo, Hyunsoo Choi, Deukjae Cho, Sekil Park, Dong-Geol Choi

专题命中 安全评测 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07221 2025-08-12 cs.LG cs.AI cs.MA stat.AP stat.ME 62%

LLM-based Agents for Automated Confounder Discovery and Subgroup Analysis in Causal Inference

Po-Han Lee, Yu-Cheng Lin, Chan-Tung Ku, Chan Hsu, Pei-Cing Huang, Ping-Hsun Wu, Yihuang Kang

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.14119 2025-08-11 cs.CL cs.AI 62%

Autonomous Structural Memory Manipulation for Large Language Models Using Hierarchical Embedding Augmentation

Derek Yotheringhay, Alistair Kirkland, Humphrey Kirkbride, Josiah Whitesteeple

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

Comments arXiv admin note: This paper has been withdrawn by arXiv due to disputed and unverifiable authorship

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.11417 2025-08-11 cs.CL cs.AI 62%

Neural Contextual Reinforcement Framework for Logical Structure Language Generation

Marcus Irvin, William Cooper, Edward Hughes, Jessica Morgan, Christopher Hamilton

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

Comments arXiv admin note: This paper has been withdrawn by arXiv due to disputed and unverifiable authorship

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02926 2025-08-08 cs.LG cs.AI cs.HC 62%

GrandJury: A Collaborative Machine Learning Model Evaluation Protocol for Dynamic Quality Rubrics

Arthur Cho

机构 * Memoirji LLC(Memiorji公司)

专题命中 安全评测 :alignment(abstract);分类 cs.AI、cs.LG

Comments 14 pages (incl. arXiv cover), 1 table, code & dataset links inside. Open-source implementation available on PyPI (grandjury package) and GitHub. Dataset available on Hugging Face under CC-BY-4.0 license

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04442 2025-08-07 cs.CL cs.AI 62%

Automated Generation of Curriculum-Aligned Multiple-Choice Questions for Malaysian Secondary Mathematics Using Generative AI

Rohaizah Abdul Wahid, Muhamad Said Nizamuddin Nadim, Suliana Sulaiman, Syahmi Akmal Shaharudin, Muhammad Danial Jupikil, Iqqwan Jasman Su Azlan Su

机构 * Fakulti Komputeran dan Meta-Teknologi (META), Universiti Pendidikan Sultan Idris(计算机与元技术学院(META),苏丹依德里斯教育大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03714 2025-08-07 cs.HC cs.AI cs.CR cs.CY 62%

"Think First, Verify Always": Training Humans to Face AI Risks

Yuksel Aydin

机构 * Independent Researcher(独立研究者)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01263 2025-08-06 cs.CL cs.AI cs.LO 62%

Bridging LLMs and Symbolic Reasoning in Educational QA Systems: Insights from the XAI Challenge at IJCNN 2025

Long S. T. Nguyen, Khang H. N. Vo, Thu H. A. Nguyen, Tuan C. Bui, Duc Q. Nguyen, Thanh-Tung Tran, Anh D. Nguyen, Minh L. Nguyen, Fabien Baldacci, Thang H. Bui, Emanuel Di Nardo, Angelo Ciaramella, Son H. Le, Ihsan Ullah, Lorenzo Di Rocco, Tho T. Quan

机构 * URA Research Group, Ho Chi Minh City University of Technology (HCMUT), Vietnam Ho Chi Minh City International University (HCMIU), Vietnam University of South-Eastern Norway, Norway Japan Advanced Institute of Science Univ. Bordeaux, CNRS, Bordeaux INP, LaBRI, UMR 5800, F-33400 Talence, France University of Naples Parthenope, Italy VNU Information Technology Institute, Vietnam National University, Vietnam Visual Intelligence Lab, School of Computer Science \& Insight Center for Data Analyitcs, University of Galway, Ireland Sapienza University of Rome, Italy

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

Comments The XAI Challenge @ TRNS-AI Workshop, IJCNN 2025: Explainable AI for Educational Question Answering. Website: https://sites.google.com/view/trns-ai/challenge/

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.14971 2025-08-06 cs.AI cs.CL cs.SD eess.AS 62%

BrainECHO: Semantic Brain Signal Decoding through Vector-Quantized Spectrogram Reconstruction for Whisper-Enhanced Text Generation

Jilong Li, Zhenxi Song, Jiaqi Wang, Meishan Zhang, Honghai Liu, Min Zhang, Zhiguo Zhang

机构 * Harbin Institute of Technology, Shenzhen, China(哈尔滨工业大学(深圳)) Peng Cheng Laboratory, China(鹏城实验室)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

Comments 8 pages (excluding references), accepted by Findings of ACL 2025

Journal ref Findings of the Association for Computational Linguistics: ACL 2025, pages 2762-2778, July 2025, Vienna, Austria

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01198 2025-08-05 cs.CL cs.AI 62%

Adaptive Content Restriction for Large Language Models via Suffix Optimization

Yige Li, Peihai Jiang, Jun Sun, Peng Shu, Tianming Liu, Zhen Xiang

机构 * Singapore Management University(新加坡管理大学) The University of Georgia(佐治亚大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

Comments 19 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.22958 2025-08-01 cs.CV cs.AI cs.LG 62%

CHECK-MAT: Checking Hand-Written Mathematical Answers for the Russian Unified State Exam

Ruslan Khrulev

机构 * Moscow State University(莫斯科国立大学)

专题命中 安全评测 :alignment(abstract);分类 cs.AI、cs.LG

Comments 15 pages, 3 figures, 10 tables. Code is available at: https://github.com/Karifannaa/Auto-check-EGE-math

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.07695 2025-08-01 cs.CL cs.AI 62%

KeyKnowledgeRAG (K^2RAG): An Enhanced RAG method for improved LLM question-answering capabilities

Hruday Markondapatnaikuni, Basem Suleiman, Abdelkarim Erradi, Shijing Chen

机构 * University of Sydney(悉尼大学) University of New South Wales(新南威尔士大学) Qatar University(卡塔尔大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

Comments 21 pages, 14 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.22576 2025-07-31 cs.CV cs.AI cs.LG 62%

COOkeD: Ensemble-based OOD detection in the era of zero-shot CLIP

Galadrielle Humblot-Renaux, Gianni Franchi, Sergio Escalera, Thomas B. Moeslund

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG

Comments accepted at ICCVW'25 - Systematic Trust in AI Models: Ensuring Fairness, Reliability, Explainability, and Accountability in Machine Learning Frameworks

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21815 2025-07-30 cs.CL cs.CY 62%

HRIPBench: Benchmarking LLMs in Harm Reduction Information Provision to Support People Who Use Drugs

Kaixuan Wang, Chenxin Diao, Jason T. Jacques, Zhongliang Guo, Shuai Zhao

机构 * University of St. Andrews(圣安德鲁大学) University of Edinburgh(爱丁堡大学) Nanyang Technological University(南洋理工大学)

专题命中 安全评测 :safety(abstract);分类 cs.CL、cs.CY

Comments 15 pages, 5 figures, 12 tables, a dataset

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21504 2025-07-30 cs.LG cs.AI 62%

Evaluation and Benchmarking of LLM Agents: A Survey

Mahmoud Mohammadi, Yipeng Li, Jane Lo, Wendy Yip

机构 * SAP Labs(SAP实验室)

专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21340 2025-07-30 cs.CL cs.AI cs.DB cs.IR 62%

StructText: A Synthetic Table-to-Text Approach for Benchmark Generation with Multi-Dimensional Evaluation

Satyananda Kashyap, Sola Shirai, Nandana Mihindukulasooriya, Horst Samulowitz

机构 * IBM Research(IBM研究院)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

Comments Data available: https://huggingface.co/datasets/ibm-research/struct-text and code available at: https://github.com/ibm/struct-text

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21188 2025-07-30 cs.LG cs.AI 62%

Embeddings to Diagnosis: Latent Fragility under Agentic Perturbations in Clinical LLMs

Raj Krishnan Vijayaraj

机构 * Independent Researcher(独立研究者)

专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.06123 2025-07-30 cs.CR cs.AI cs.CV cs.LG 62%

Adversarial attacks and defenses in explainable artificial intelligence: A survey

Hubert Baniecki, Przemyslaw Biecek

机构 * University of Warsaw(华沙大学) Warsaw University of Technology(华沙理工大学)

专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.LG

Comments Accepted by Information Fusion

Journal ref Information Fusion, vol. 107, 102303, 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20280 2025-07-29 cs.AI cs.CL 62%

SciToolAgent: A Knowledge Graph-Driven Scientific Agent for Multi-Tool Integration

Keyan Ding, Jing Yu, Junjie Huang, Yuchen Yang, Qiang Zhang, Huajun Chen

机构 * College of Computer Science and Technology(计算机科学与技术学院) Zhejiang University(浙江大学) Zhejiang Key Laboratory of Intelligent Manufacturing for Functional Chemicals(功能化学品智能制造重点实验室) ZJU-Hangzhou Global Scientific and Technological Innovation Center(浙大杭州国际科学技术创新中心) ZJU-UIUC Institute(浙大UIUC研究院) The Polytechnic Institute(技术学院) State Key Laboratory of Ocean Sensing(海洋感知国家重点实验室)

专题命中 安全评测 :safety(abstract);分类 cs.CL、cs.AI

Comments 21 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.16802 2025-07-29 cs.CL cs.LG 62%

Agentar-Fin-R1: Enhancing Financial Intelligence through Domain Expertise, Training Efficiency, and Advanced Reasoning

Yanjun Zheng, Xiyang Du, Longfei Liao, Xiaoke Zhao, Zhaowen Zhou, Jingze Song, Bo Zhang, Jiawei Liu, Xiang Qi, Zhe Li, Zhiqiang Zhang, Wei Wang, Peng Zhang

机构 * Ant Group(蚂蚁集团)

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19803 2025-07-29 cs.LG cs.AI 62%

AI-Based Clinical Rule Discovery for NMIBC Recurrence through Tsetlin Machines

Saram Abbas, Naeem Soomro, Rishad Shafik, Rakesh Heer, Kabita Adhikari

机构 * Newcastle University, UK Freeman Hospital, UK Imperial College London \& Newcastle University Centre for Care School of Engineering Newcastle University, UK

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG

Comments Submitted to ISTM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19556 2025-07-29 cs.CY cs.AI 62%

PEMUTA: Pedagogically-Enriched Multi-Granular Undergraduate Thesis Assessment

Jialu Zhang, Qingyang Sun, Qianyi Wang, Weiyi Zhang, Zunjie Xiao, Xiaoqing Zhang, Jianfeng Ren, Jiang Liu

机构 * Research Institute of Trustworthy Autonomous Systems and Department of Computer Science and Engineering(可信自主系统研究 institute 和计算机科学与工程系) Southern University of Science and Technology(南方科技大学) School of Computer Science, University of Nottingham Ningbo China(宁波大学计算机学院) School of Ophthalmology and Optometry, Wenzhou Medical University(温州医学院眼视光学院)

专题命中 安全评测 :alignment(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.15956 2025-07-29 cs.CL cs.AI 62%

Do Large Language Models Have an English Accent? Evaluating and Improving the Naturalness of Multilingual LLMs

Yanzhu Guo, Simone Conia, Zelin Zhou, Min Li, Saloni Potdar, Henry Xiao

机构 * Apple(苹果公司) Inria Paris(巴黎国家信息与自动化研究所) École Polytechnique(巴黎高等理工学院) Sapienza University of Rome(罗马萨皮恩扎大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

Comments ACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19132 2025-07-28 cs.AI cs.CL cs.CV cs.HC 62%

OS-MAP: How Far Can Computer-Using Agents Go in Breadth and Depth?

Xuetian Chen, Yinghao Chen, Xinfeng Yuan, Zhuo Peng, Lu Chen, Yuekeng Li, Zhoujia Zhang, Yingqian Huang, Leyan Huang, Jiaqing Liang, Tianbao Xie, Zhiyong Wu, Qiushi Sun, Biqing Qi, Bowen Zhou

机构 * Fudan University(复旦大学) Shanghai AI Lab(上海人工智能实验室) Tsinghua University(清华大学) The University of Hong Kong(香港大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

Comments Work in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.18918 2025-07-28 cs.CL cs.AI 62%

Uncovering Cross-Linguistic Disparities in LLMs using Sparse Autoencoders

Richmond Sin Jing Xuan, Jalil Huseynov, Yang Zhang

机构 * National University of Singapore(新加坡国立大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏