arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-10-24 至 2025-10-24 共收录 60 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 24 篇

2510.20039 2025-10-24 cs.HC cs.AI cs.CL cs.CY 67%

Beyond One-Way Influence: Bidirectional Opinion Dynamics in Multi-Turn Human-LLM Interactions

Yuyang Jiang, Longjie Guo, Yuchen Wu, Aylin Caliskan, Tanu Mitra, Hua Shen

机构 * University of Chicago(芝加哥大学) New York University(纽约大学) University of Washington(华盛顿大学) New York University Shanghai(纽约大学上海)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI、cs.CY

Comments 26 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20611 2025-10-24 cs.LG cs.AI 62%

PSO-XAI: A PSO-Enhanced Explainable AI Framework for Reliable Breast Cancer Detection

Mirza Raquib, Niloy Das, Farida Siddiqi Prity, Arafath Al Fahim, Saydul Akbar Murad, Mohammad Amzad Hossain, MD Jiabul Hoque, Mohammad Ali Moni

机构 * Department of Information and Communication Engineering, Noakhali Science and Technology University(信息与通信工程系,诺阿克利科学与技术大学) Department of Computer Science & Engineering, Netrokona University(计算机科学与工程系,纳特罗克纳大学) Department of Computer and Communication Engineering, International Islamic University Chittagong(计算机与通信工程系,国际伊斯兰大学查塔格ONG) Department of Mechatronics and Industrial Engineering, Chittagong University of Engineering and Technology(机电与工业工程系,查塔格ONG工程与技术大学) School of Computing Sciences and Computer Engineering, University of Southern Mississippi(计算科学与计算机工程学院,密西西比州立大学) AI & Digital Health Technology, Artificial Intelligence and Cyber Futures Institute, Charles Sturt University(人工智能与数字健康技术,人工智能与网络未来研究院,查尔斯·斯图尔特大学) AI & Digital Health Technology, Rural Health Research Institute, Charles Sturt University(人工智能与数字健康技术,农村健康研究学院,查尔斯·斯图尔特大学)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20270 2025-10-24 cs.LG cs.CL 62%

ImpossibleBench: Measuring LLMs' Propensity of Exploiting Test Cases

Ziqian Zhong, Aditi Raghunathan, Nicholas Carlini

专题命中 安全评测 :safety(abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20211 2025-10-24 cs.SE cs.AI cs.LG 62%

Automated Cloud Infrastructure-as-Code Reconciliation with AI Agents

Zhenning Yang, Hui Guan, Victor Nicolet, Brandon Paulsen, Joey Dodds, Daniel Kroening, Ang Chen

机构 * University of Michigan(密歇根大学) Amazon Web Services(亚马逊网络服务)

专题命中 安全评测 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20193 2025-10-24 cs.IR cs.CL cs.CV cs.LG 62%

Multimedia-Aware Question Answering: A Review of Retrieval and Cross-Modal Reasoning Architectures

Rahul Raja, Arpita Vats

机构 * Carnegie Mellon University(卡内基梅隆大学) Boston University(波士顿大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.LG

Comments In Proceedings of the 2nd ACM Workshop in AI-powered Question and Answering Systems (AIQAM '25), October 27-28, 2025, Dublin, Ireland. ACM, New York, NY, USA, 8 pages. https://doi.org/10.1145/3746274.3760393

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11824 2025-10-24 cs.MA cs.AI cs.LG 62%

Empirical Study on Robustness and Resilience in Cooperative Multi-Agent Reinforcement Learning

Simin Li, Zihao Mao, Hanxiao Li, Zonglei Jing, Zhuohang bian, Jun Guo, Li Wang, Zhuoran Han, Ruixiao Xu, Xin Yu, Chengdong Ma, Yuqing Ma, Bo An, Yaodong Yang, Weifeng Lv, Xianglong Liu

机构 * State Key Laboratory of Complex & Critical Software Environment(复杂与关键软件环境国家重点实验室) Beihang University(北京航空航天大学) Zhongguancun Laboratory(中关村实验室) Institute of Artificial Intelligence, Peking University(北京大学人工智能研究院) Institute of data space, Hefei Comprehensive National Science Center(合肥综合国家科学中心数据空间研究院) Nanyang Technological University(南洋理工大学)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG

Comments 44 pages, 16 figures, NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.07142 2025-10-24 cs.CL cs.AI cs.DL 62%

Toward Purpose-oriented Topic Model Evaluation enabled by Large Language Models

Zhiyin Tan, Jennifer D'Souza

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

Comments Accepted for publication in International Journal on Digital Libraries (IJDL)

Journal ref International Journal on Digital Libraries, vol. 26, no. 4, pp. 23, December 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.11261 2025-10-24 cs.AI cs.CL 62%

Sycophancy in Vision-Language Models: A Systematic Analysis and an Inference-Time Mitigation Framework

Yunpu Zhao, Rui Zhang, Junbin Xiao, Changxin Ke, Ruibo Hou, Yifan Hao, Ling Li

机构 * School of Computer Science and Technology, University of Science and Technology of China(计算机科学与技术学院,中国科学技术大学) State Key Lab of Processors, Institute of Computing Technology, Chinese Academy of Sciences(处理器国家重点实验室,中国科学院计算技术研究所) Department of Computer Science, National University of Singapore(计算机科学系,新加坡国立大学) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Intelligent Software Research Center, Institute of Software, Chinese Academy of Sciences(软件智能研究中心,中国科学院软件研究所)

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL、cs.AI

Journal ref Neurocomputing, Volume 659, 2026, 131217

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.19866 2025-10-24 cs.CL cs.AI 62%

An Evaluation of the Pedagogical Soundness and Usability of AI-Generated Lesson Plans Across Different Models and Prompt Frameworks in High-School Physics

Xincheng Liu

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

Comments 20 pages, 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.20612 2025-10-24 cs.CV cs.CL cs.LG 62%

Roboflow100-VL: A Multi-Domain Object Detection Benchmark for Vision-Language Models

Peter Robicheaux, Matvei Popov, Anish Madan, Isaac Robinson, Joseph Nelson, Deva Ramanan, Neehar Peri

机构 * Roboflow Carnegie Mellon University(卡内基梅隆大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.LG

Comments The first two authors contributed equally. This work has been accepted to the Neural Information Processing Systems (NeurIPS) 2025 Datasets & Benchmark Track. Project Page: https://rf100-vl.org/

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20699 2025-10-24 q-fin.CP cs.AI 57%

Fusing Narrative Semantics for Financial Volatility Forecasting

Yaxuan Kong, Yoontae Hwang, Marcus Kaiser, Chris Vryonides, Roel Oomen, Stefan Zohren

机构 * University of Oxford(牛津大学) Pusan National University(釜山国立大学)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

Comments The 6th ACM International Conference on AI in Finance (ICAIF 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20337 2025-10-24 cs.AI 57%

Collateral Damage Assessment Model for AI System Target Engagement in Military Operations

Clara Maathuis, Kasper Cools

机构 * Open University of the Netherlands(荷兰开放大学) Royal Military Academy, Belgium(比利时皇家军事学院) Vrije Universiteit Brussel, Belgium(比利时自由大学)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

Comments Accepted at MILCOM 2025 WS07

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11145 2025-10-24 cs.CL cs.PL 57%

Text2Mem: A Unified Memory Operation Language for Memory Operating System

Yi Wang, Lihai Yang, Boyu Chen, Gongyi Zou, Kerun Xu, Bo Tang, Feiyu Xiong, Siheng Chen, Zhiyu Li

机构 * MemTensor (Shanghai) Technology(MemTensor(上海)技术有限公司) University of Oxford(牛津大学) National University of Singapore(新加坡国立大学) Shanghai Jiao Tong University(上海交通大学)

专题命中 安全评测 :safety(abstract);分类 cs.CL

Comments 12 pages, 3 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.02464 2025-10-24 cs.CL 57%

SpecEval: Evaluating Model Adherence to Behavior Specifications

Ahmed Ahmed, Kevin Klyman, Yi Zeng, Sanmi Koyejo, Percy Liang

机构 * Department of Computer Science, Stanford University(计算机科学系,斯坦福大学) The Bradley Department of Electrical and Computer Engineering, Virginia Tech(电气与计算机工程系,弗吉尼亚理工学院)

专题命中 安全评测 :safety(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20352 2025-10-24 cs.CL 57%

RMTBench: Benchmarking LLMs Through Multi-Turn User-Centric Role-Playing

Hao Xiang, Tianyi Tang, Yang Su, Bowen Yu, An Yang, Fei Huang, Yichang Zhang, Yaojie Lu, Hongyu Lin, Xianpei Han, Jingren Zhou, Junyang Lin, Le Sun

机构 * Chinese Information Processing Laboratory, Institute of Software, Chinese Academy of Sciences(中国科学院软件研究所信息处理实验室) Alibaba Group(阿里巴巴集团) University of Chinese Academy of Sciences(中国科学院大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19271 2025-10-24 cs.SE cs.AI cs.PL 57%

Fine-Tuning Multilingual Language Models for Code Review: An Empirical Study on Industrial C# Projects

Igli Begolli, Meltem Aksoy, Daniel Neider

机构 * Technical University Dortmund, Lovion GmbH(图鲁尼大学,洛维昂公司) Security\ Alliance Ruhr, Technical University Dortmund(安全联盟鲁尔,图鲁尼大学)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.16247 2025-10-24 cs.CV cs.LG 57%

OpenMIBOOD: Open Medical Imaging Benchmarks for Out-Of-Distribution Detection

Max Gutbrod, David Rauber, Danilo Weber Nunes, Christoph Palm

专题命中 安全评测 :trustworthy(abstract);分类 cs.LG

Comments Updated results for NNGuide and ViM

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.19836 2025-10-24 cs.AI cs.SY eess.SY 57%

Benchmarking Reasoning Reliability in Artificial Intelligence Models for Energy-System Analysis

Eliseo Curcio

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20284 2025-10-24 cs.CV 50%

Knowledge-Informed Neural Network for Complex-Valued SAR Image Recognition

Haodong Yang, Zhongling Huang, Shaojie Guo, Zhe Zhang, Gong Cheng, Junwei Han

机构 * School of Automation, Northwestern Polytechnical University(西北工业大学自动化学院) Shenzhen Research Institute of Northwestern Polytechnical University(西北工业大学深圳研究院) School of Artificial Intelligence, Chongqing University of Posts and Telecommunications(重庆邮电大学人工智能学院) Aerospace Information Technology University(航天信息大学) Suzhou Aerospace Information Research Institute(苏州航天信息研究所) National Key Laboratory of Microwave Imaging(微波成像国家重点实验室) Aerospace Information Research Institute, CAS(中国科学院航天信息研究所) School of Electronic, Electrical and Communication Engineering, University of Chinese Academy of Sciences(中国科学院大学电子、电气与通信工程学院)

专题命中 安全评测 :trustworthy(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

2. AI治理与伦理 3 篇

2505.09131 2025-10-24 cs.LG cs.AI stat.ML 81%

Fair Clustering via Alignment

Kunwoong Kim, Jihu Lee, Sangchul Park, Yongdai Kim

机构 * Department of Statistics, Seoul National University, Republic of Korea(统计系,首尔国立大学,大韩民国) School of Law, Seoul National University, Republic of Korea(法学院,首尔国立大学,大韩民国)

专题命中 AI治理与伦理 :alignment(title,abstract);分类 cs.AI、cs.LG

Journal ref ICML 2025 (Forty-Second International Conference on Machine Learning)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20782 2025-10-24 cs.CL cs.AI 62%

A Use-Case Specific Dataset for Measuring Dimensions of Responsible Performance in LLM-generated Text

Alicia Sagae, Chia-Jung Lee, Sandeep Avula, Brandon Dang, Vanessa Murdock

机构 * AWS Responsible AI(AWS负责任人工智能)

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CL、cs.AI

Comments 24 pages with 3 figures, to appear in Proceedings of the 34th ACM International Conference on Information and Knowledge Management (CIKM '25)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.19733 2025-10-24 cs.CL cs.LG 62%

Zhyper: Factorized Hypernetworks for Conditioned LLM Fine-Tuning

M. H. I. Abdalla, Zhipin Wang, Christian Frey, Steffen Eger, Josif Grabocka

机构 * Department of Computer Science University of Technology Nuremberg(计算机科学系图腾技术大学纽伦堡)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 其他安全 8 篇

2510.20377 2025-10-24 cs.AI cs.CL 62%

IKnow: Instruction-Knowledge-Aware Continual Pretraining for Effective Domain Adaptation

Tianyi Zhang, Florian Mai, Lucie Flek

机构 * University of Bonn(波恩大学) Lamarr Institute for Machine Learning and Artificial Intelligence(拉玛尔机器学习与人工智能研究所)

专题命中 其他安全 :safety(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.16722 2025-10-24 cs.CL cs.AI 62%

Breaking mBad! Supervised Fine-tuning for Cross-Lingual Detoxification

Himanshu Beniwal, Youngwoo Kim, Maarten Sap, Soham Dan, Thomas Hartvigsen

机构 * Indian Institute of Technology Gandhinagar(印度古吉拉特邦理工学院加尔文加尔)

专题命中 其他安全 :safety(abstract);分类 cs.CL、cs.AI

Comments Accepted at MELT Workshop @ COLM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20428 2025-10-24 cs.LG 57%

An Empirical Study of Sample Selection Strategies for Large Language Model Repair

Xuran Li, Jingyi Wang

机构 * Zhejiang University(浙江大学)

专题命中 其他安全 :safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20244 2025-10-24 cs.CV cs.LG 57%

Empower Words: DualGround for Structured Phrase and Sentence-Level Temporal Grounding

Minseok Kang, Minhyeok Lee, Minjung Kim, Donghyeong Kim, Sangyoun Lee

机构 * Yonsei University(延世大学) LG Electronics(LG电子)

专题命中 其他安全 :alignment(abstract);分类 cs.LG

Comments Comments: 28 pages, including appendix. 5 figures. Full version of the NeurIPS 2025 paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25033 2025-10-24 cs.CV cs.LG 57%

VT-FSL: Bridging Vision and Text with LLMs for Few-Shot Learning

Wenhao Li, Qiangchang Wang, Xianjing Meng, Zhibin Wu, Yilong Yin

机构 * School of Software, Shandong University(山东大学软件学院) Shenzhen Loop Area Institute(深圳河套学院) School of Computing and Artificial Intelligence, Shandong University of Finance and Economics(山东财经大学计算机与人工智能学院)

专题命中 其他安全 :alignment(abstract);分类 cs.LG

Comments Accepted by NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.19529 2025-10-24 cs.CL 57%

Blockwise SFT for Diffusion Language Models: Reconciling Bidirectional Attention and Autoregressive Decoding

Bowen Sun, Yujun Cai, Ming-Hsuan Yang, Yiwei Wang

机构 * University of California, Merced(加州大学默塞德分校) The University of Queensland(昆士兰大学) Google DeepMind(谷歌DeepMind)

专题命中 其他安全 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.19201 2025-10-24 cs.CL 57%

DREAM: Drafting with Refined Target Features and Entropy-Adaptive Cross-Attention Fusion for Multimodal Speculative Decoding

Yunhai Hu, Tianhua Xia, Zining Liu, Rahul Raman, Xingyu Liu, Bo Bao, Eric Sather, Vithursan Thangarasa, Sai Qian Zhang

机构 * Courant Institute of Mathematical Sciences, New York University(纽约大学数学科学学院) Tandon School of Engineering, New York University(纽约大学工程学院) Cerebras Systems Inc.(Cerebras Systems公司) University of Pennsylvania(宾夕法尼亚大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.14349 2025-10-24 cs.CV 50%

Vision-Centric Activation and Coordination for Multimodal Large Language Models

Yunnan Wang, Fan Lu, Kecheng Zheng, Ziyuan Huang, Ziqiang Li, Wenjun Zeng, Xin Jin

机构 * MoE Key Lab of Artificial Intelligence, Shanghai Jiao Tong University(人工智能MoE实验室,上海交通大学) Ant Group(蚂蚁集团) Ningbo Institute of Digital Twin, Eastern Institute of Technology, Ningbo(宁波数字孪生研究所,东部技术研究所,宁波)

专题命中 其他安全 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏