arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 9434 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 9434 篇

2509.21143 2025-09-30 cs.RO cs.CL 57%

Automotive-ENV: Benchmarking Multimodal Agents in Vehicle Interface Systems

Junfeng Yan, Biao Wu, Meng Fang, Ling Chen

机构 * Australian Artificial Intelligence Institute(澳大利亚人工智能研究所) University of Liverpool(利物浦大学)

专题命中 安全评测 :safety(abstract);分类 cs.CL

Comments 10 pages, 5 figures,

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.10036 2025-09-30 cs.CL 57%

DataPuzzle: Breaking Free from the Hallucinated Promise of LLMs in Data Analysis

Zhengxuan Zhang, Zhuowen Liang, Yin Wu, Teng Lin, Yuyu Luo, Nan Tang

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21949 2025-09-29 cs.NI cs.CL 57%

Evaluating Open-Source Large Language Models for Technical Telecom Question Answering

Arina Caraus, Alessio Buscemi, Sumit Kumar, Ion Turcanu

机构 * Luxembourg Institute of Science and Technology (LIST)(卢森堡科学与技术研究院) RMT Labs(RMT实验室)

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL

Comments Accepted at the IEEE GLOBECOM Workshops 2025: "Large AI Model over Future Wireless Networks"

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21749 2025-09-29 cs.CL cs.SD 57%

Thinking with Sound: Audio Chain-of-Thought Enables Multimodal Reasoning in Large Audio-Language Models

Zhen Xiong, Yujun Cai, Zhecheng Li, Junsong Yuan, Yiwei Wang

机构 * University of Southern California(南加州大学) University of Queensland(昆士兰大学) University of California, San Diego(加州大学圣地亚哥分校) University of Buffalo(布法罗大学) University of California, Merced(加州大学默塞德分校)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21662 2025-09-29 cs.LG 57%

MMPlanner: Zero-Shot Multimodal Procedural Planning with Chain-of-Thought Object State Reasoning

Afrina Tabassum, Bin Guo, Xiyao Ma, Hoda Eldardiry, Ismini Lourentzou

机构 * Amazon(亚马逊公司) Alexa, Amazon(亚马逊Alexa部门) Virginia Tech(弗吉尼亚理工大学) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

专题命中 安全评测 :alignment(abstract);分类 cs.LG

Comments 17 pages, 9 figures, 14 tables, Findings of the Association for Computational Linguistics: EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21450 2025-09-29 cs.CL 57%

LLM-Based Support for Diabetes Diagnosis: Opportunities, Scenarios, and Challenges with GPT-5

Gaurav Kumar Gupta, Nirajan Acharya, Pranal Pande

机构 * Youngstown State University(扬斯敦州立大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17978 2025-09-29 cs.AI cs.LO 57%

The STAR-XAI Protocol: A Framework for Inducing and Verifying Agency, Reasoning, and Reliability in AI Agents

Antoni Guasch, Maria Isabel Valdez

机构 * Ixent Games

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

Comments Version 2: This article consolidates and replaces a previous version to present the complete research in a single, comprehensive manuscript

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21318 2025-09-26 cs.CV cs.AI 57%

SD3.5-Flash: Distribution-Guided Distillation of Generative Flows

Hmrishav Bandyopadhyay, Rahim Entezari, Jim Scott, Reshinth Adithyan, Yi-Zhe Song, Varun Jampani

机构 * Stability AI SketchX, University of Surrey(SketchX,大学)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

Comments Project Page: https://hmrishavbandy.github.io/sd35flash/

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21310 2025-09-26 cs.AI 57%

SAGE: A Realistic Benchmark for Semantic Understanding

Samarth Goel, Reagan J. Lee, Kannan Ramchandran

机构 * University of California, Berkeley(加州大学伯克利分校)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

Comments 39th Conference on Neural Information Processing Systems (NeurIPS 2025) Workshop: Evaluating the Evolving LLM Lifecycle: Benchmarks, Emergent Abilities, and Scaling

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21208 2025-09-26 cs.CL 57%

CLaw: Benchmarking Chinese Legal Knowledge in Large Language Models - A Fine-grained Corpus and Reasoning Analysis

Xinzhe Xu, Liang Zhao, Hongshen Xu, Chen Chen

机构 * Peking University(北京大学) LLM-Core

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.20418 2025-09-26 cs.CR cs.AI cs.ET 57%

A Taxonomy of Data Risks in AI and Quantum Computing (QAI) - A Systematic Review

Grace Billiris, Asif Gill, Madhushi Bandara

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

Comments 11 pages, 2 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.08810 2025-09-25 cs.CR cs.AI 57%

Machine Learning-Based Detection of DDoS Attacks in VANETs for Emergency Vehicle Communication

Bappa Muktar, Vincent Fono, Adama Nouboukpo

机构 * Department of Computer Science University of Quebec in Outaouais (UQO)(计算机科学系魁北克大学 Outaouais 分校)

专题命中 安全评测 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19604 2025-09-25 cs.LG 57%

Improved Therapeutic Antibody Reformatting through Multimodal Machine Learning

Jiayi Xin, Aniruddh Raghu, Nick Bhattacharya, Adam Carr, Melanie Montgomery, Hunter Elliott

机构 * University of Pennsylvania(宾夕法尼亚大学) BigHat Biosciences(BigHat生物技术公司)

专题命中 安全评测 :safety(abstract);分类 cs.LG

Comments NeurIPS 2025 AI4Science Workshop and NeurIPS 2025 Multi-modal Foundation Models and Large Language Models for Life Sciences Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.01301 2025-09-25 cs.CL 57%

Culture is Everywhere: A Call for Intentionally Cultural Evaluation

Juhyun Oh, Inha Cha, Michael Saxon, Hyunseung Lim, Shaily Bhatt, Alice Oh

专题命中 安全评测 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.00935 2025-09-25 cs.CR cs.AI 57%

Measuring Harmfulness of Computer-Using Agents

Aaron Xuxiang Tian, Ruofan Zhang, Janet Tang, Ji Wang, Tianyu Shi, Jiaxin Wen

机构 * Arizona State University(亚利桑那州立大学) University of California, Berkeley(加州大学伯克利分校)

专题命中 安全评测 :safety(abstract);分类 cs.AI

Comments 17 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23368 2025-09-25 cs.CL 57%

Threading the Needle: Reweaving Chain-of-Thought Reasoning to Explain Human Label Variation

Beiduo Chen, Yang Janet Liu, Anna Korhonen, Barbara Plank

专题命中 安全评测 :alignment(abstract);分类 cs.CL

Comments Accepted by EMNLP 2025 Main (Oral), 25 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.16180 2025-09-25 cs.CV cs.CL 57%

Redemption Score: A Multi-Modal Evaluation Framework for Image Captioning via Distributional, Perceptual, and Linguistic Signal Triangulation

Ashim Dahal, Ankit Ghimire, Saydul Akbar Murad, Nick Rahimi

机构 * University of Southern Mississippi(密苏里州南方大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.00046 2025-09-24 cs.CV cs.AI 57%

Leveraging Large Models to Evaluate Novel Content: A Case Study on Advertisement Creativity

Zhaoyi Joey Hou, Adriana Kovashka, Xiang Lorraine Li

机构 * Department of Computer Science University of Pittsburgh(计算机科学系宾夕法尼亚大学)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

Comments To Appear in EMNLP2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18568 2025-09-24 cs.LG 57%

Explainable Graph Neural Networks: Understanding Brain Connectivity and Biomarkers in Dementia

Niharika Tewari, Nguyen Linh Dan Le, Mujie Liu, Jing Ren, Ziqi Xu, Tabinda Sarwar, Veeky Baths, Feng Xia

机构 * School of Computing Technologies RMIT University Melbourne VIC Australia Department of Biological Sciences Department of Computer Science \& Information Systems Birla Institute of Technology Institute of Innovation, Science Sustainability Federation University Australia Ballarat VIC Australia RMIT University Birla Institute of Technology Federation University Australia

专题命中 安全评测 :trustworthy(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.14498 2025-09-23 cs.CL 57%

LLaSA: A Sensor-Aware LLM for Natural Language Reasoning of Human Activity from IMU Data

Sheikh Asif Imran, Mohammad Nur Hossain Khan, Subrata Biswas, Bashima Islam

机构 * Worcester Polytechnic Institute(沃斯特理工大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17249 2025-09-23 cs.CL 57%

Extending Automatic Machine Translation Evaluation to Book-Length Documents

Kuang-Da Wang, Shuoyang Ding, Chao-Han Huck Yang, Ping-Chun Hsieh, Wen-Chih Peng, Vitaly Lavrukhin, Boris Ginsburg

机构 * National Yang Ming Chiao Tung University(国家阳明交通大学) NVIDIA

专题命中 安全评测 :alignment(abstract);分类 cs.CL

Comments Accepted for EMNLP 2025 main conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17343 2025-09-23 cs.SE cs.AI 57%

Agentic AI for Software: thoughts from Software Engineering community

Abhik Roychoudhury

机构 * NUS(国立新加坡大学) SonarSource SA

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

Comments 4 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.11023 2025-09-23 cs.LG 57%

Informed, but Not Always Improved: Challenging the Benefit of Background Knowledge in GNNs

Kutalmış Coşkun, Ivo Kavisanczki, Amin Mirzaei, Tom Siegl, Bjarne C. Hiller, Stefan Lüdtke, Martin Becker

机构 * University of Rostock(罗斯托克大学) University of Rostock, University of Marburg(罗斯托克大学、马堡大学)

专题命中 安全评测 :alignment(abstract);分类 cs.LG

Comments 10 pages, 7 figures, added repo link

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.01830 2025-09-23 cs.CL 57%

From Language to Cognition: How LLMs Outgrow the Human Language Network

Badr AlKhamissi, Greta Tuckute, Yingtian Tang, Taha Binhuraib, Antoine Bosselut, Martin Schrimpf

机构 * EPFL(苏黎世联邦理工学院) MIT(麻省理工学院) Georgia Institute of Technology(佐治亚理工学院)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

Comments EMNLP 2025. Project Page at https://language-to-cognition.epfl.ch

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.01609 2025-09-23 cs.CL 57%

Adaptive Distraction: Probing LLM Contextual Robustness with Automated Tree Search

Yanbo Wang, Zixiang Xu, Yue Huang, Chujie Gao, Siyuan Wu, Jiayi Ye, Pin-Yu Chen, Xiuying Chen, Xiangliang Zhang

机构 * Mohamed bin Zayed University of Artificial Intelligence (MBZUAI)(莫扎伊德大学人工智能学院) University of Notre Dame(诺特丹大学) IBM Research(IBM研究院)

专题命中 安全评测 :DPO(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16543 2025-09-23 cs.CL 57%

ChemOrch: Empowering LLMs with Chemical Intelligence via Synthetic Instructions

Yue Huang, Zhengzhe Jiang, Xiaonan Luo, Kehan Guo, Haomin Zhuang, Yujun Zhou, Zhengqing Yuan, Xiaoqi Sun, Jules Schleinitz, Yanbo Wang, Shuhao Zhang, Mihir Surve, Nitesh V Chawla, Olaf Wiest, Xiangliang Zhang

机构 * Department of Computer Science and Engineering, University of Notre Dame(诺丁汉大学计算机科学与工程系) MIT(麻省理工学院) CalTech(加州理工学院) MBZUAI(穆斯林人工智能研究所) CMU(卡内基梅隆大学) Department of Chemistry & Biochemistry, University of Notre Dame(诺丁汉大学化学与生物化学系)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16275 2025-09-23 cs.CR cs.AI cs.SE 57%

SecureFixAgent: A Hybrid LLM Agent for Automated Python Static Vulnerability Repair

Jugal Gajjar, Kamalasankari Subramaniakuppusamy, Relsy Puthal, Kaustik Ranaware

机构 * Computer Science Department(计算机科学系) Applied Economics Department(应用经济学系)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

Comments 6 pages, 3 figures, 4 tables, 1 algorithm, accepted in the Robustness and Security of Large Language Models (ROSE-LLM) special session at ICMLA 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16268 2025-09-23 cs.SE cs.AI 57%

Digging Into the Internal: Causality-Based Analysis of LLM Function Calling

Zhenlan Ji, Daoyuan Wu, Wenxuan Wang, Pingchuan Ma, Shuai Wang, Lei Ma

专题命中 安全评测 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.18831 2025-09-23 cs.CL 57%

Measuring Risk of Bias in Biomedical Reports: The RoBBR Benchmark

Jianyou Wang, Weili Cao, Longtian Bao, Youze Zheng, Gil Pasternak, Kaicheng Wang, Xiaoyue Wang, Ramamohan Paturi, Leon Bergen

机构 * Laboratory for Emerging Intelligence(新兴智能实验室) University of California, San Diego(加州大学圣地亚哥分校)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

Comments Published at EMNLP 2025 (Main)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15239 2025-09-22 cs.CL 57%

WangchanThaiInstruct: An instruction-following Dataset for Culture-Aware, Multitask, and Multi-domain Evaluation in Thai

Peerat Limkonchotiwat, Pume Tuchinda, Lalita Lowphansirikul, Surapon Nonesung, Panuthep Tasawong, Alham Fikri Aji, Can Udomcharoenchaikit, Sarana Nutanong

机构 * AI Singapore(AI新加坡) VISTEC MBZUAI

专题命中 安全评测 :alignment(abstract);分类 cs.CL

Comments Accepted to EMNLP 2025 (Main). Model and Dataset: https://huggingface.co/collections/airesearch/wangchan-thai-instruction-6835722a30b98e01598984fd

详情

展开后加载摘要…

URL PDF HTML 收藏