arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-10-28 至 2025-10-28 共收录 84 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 越狱攻击 5 篇

2505.19911 2025-10-28 cs.CV 50%

Attention! Your Vision Language Model Could Be Maliciously Manipulated

Xiaosen Wang, Shaokang Wang, Zhijin Ge, Yuyang Luo, Shudong Zhang

专题命中 越狱攻击 :trustworthy(abstract)

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 幻觉与事实性 8 篇

2306.11593 2025-10-28 cs.CV cs.AI cs.CL cs.DB cs.LG 67%

Improving Image Captioning Descriptiveness by Ranking and LLM-based Fusion

Luigi Celona, Simone Bianco, Marco Donzella, Paolo Napoletano

机构 * Department of Informatics, Systems and Communication(信息学、系统与通信系)

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments This manuscript has been accepted for publication in Springer Neural Computing and Applications

Journal ref Neural Computer & Application 37, 27279-27299 (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23264 2025-10-28 cs.LG cs.AI 62%

PAHQ: Accelerating Automated Circuit Discovery through Mixed-Precision Inference Optimization

Xinhai Wang, Shu Yang, Liangyu Wang, Lin Zhang, Huanyi Xie, Lijie Hu, Di Wang

机构 * King Abdullah University of Science and Technology(卡布斯大学) Harbin Institute of Technology(哈尔滨工业大学)

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22261 2025-10-28 cs.LG cs.AI 62%

Epistemic Deep Learning: Enabling Machine Learning Models to Know When They Do Not Know

Shireen Kudukkil Manchingal

专题命中 幻觉与事实性 :safety(abstract);分类 cs.AI、cs.LG

Comments PhD thesis

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22751 2025-10-28 cs.AI cs.CL 62%

Multi-Modal Fact-Verification Framework for Reducing Hallucinations in Large Language Models

Piyushkumar Patel

机构 * Microsoft(微软)

专题命中 幻觉与事实性 :trustworthy(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22362 2025-10-28 cs.LG cs.CL 62%

Mapping Faithful Reasoning in Language Models

Jiazheng Li, Andreas Damianou, J Rosser, José Luis Redondo García, Konstantina Palla

机构 * King’s College London(伦敦国王学院) Spotify UK(Spotify英国分公司) University of Oxford(牛津大学) Spotify Spain(Spotify西班牙分公司)

专题命中 幻觉与事实性 :safety(abstract);分类 cs.CL、cs.LG

Comments 9 pages, Accepted to the Mechanistic Interpretability Workshop at NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22614 2025-10-28 cs.SE cs.AI 57%

Does In-IDE Calibration of Large Language Models work at Scale?

Roham Koohestani, Agnia Sergeyuk, David Gros, Claudio Spiess, Sergey Titov, Prem Devanbu, Maliheh Izadi

机构 * Delft University of Technology(代尔夫特理工大学) JetBrains Research(JetBrains研究) University of California, Davis(加州大学戴维斯分校)

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.AI

Comments Under Review

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.16503 2025-10-28 cs.CV 50%

Noise Diffusion for Enhancing Semantic Faithfulness in Text-to-Image Synthesis

Boming Miao, Chunxiao Li, Xiaoxiao Wang, Andi Zhang, Rui Sun, Zizhe Wang, Yao Zhu

机构 * Beijing Normal University(北京师范大学) University of Chinese Academy of Sciences(中国科学院大学) University of Manchester(曼彻斯特大学) The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)) Tsinghua University(清华大学)

专题命中 幻觉与事实性 :alignment(abstract)

Comments Updated author formatting; no substantive changes

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.07759 2025-10-28 cs.IR 50%

A Survey of Long-Document Retrieval in the PLM and LLM Era

Minghan Li, Miyang Luo, Tianrui Lv, Yishuai Zhang, Siqi Zhao, Ercong Nie, Guodong Zhou

专题命中 幻觉与事实性 :alignment(abstract)

Comments 32 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 隐私与版权 3 篇

2510.22070 2025-10-28 cs.LG cs.CV eess.IV stat.ML 70%

MAGIC-Flow: Multiscale Adaptive Conditional Flows for Generation and Interpretable Classification

Luca Caldera, Giacomo Bottacini, Lara Cavinato

机构 * MOX, Department of Mathematics Politecnico di Milano(米兰理工大学数学系MOX部门)

专题命中 隐私与版权 :alignment(abstract);trustworthy(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23019 2025-10-28 cs.LG cs.DC 57%

Sentinel: Dynamic Knowledge Distillation for Personalized Federated Intrusion Detection in Heterogeneous IoT Networks

Gurpreet Singh, Keshav Sood, P. Rajalakshmi, Yong Xiang

专题命中 隐私与版权 :alignment(abstract);分类 cs.LG

Comments This is a preprint version of a paper currently under review for possible publication in IEEE TDSC

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22964 2025-10-28 cs.CV 50%

Survey of Multimodal Geospatial Foundation Models: Techniques, Applications, and Challenges

Liling Yang, Ning Chen, Jun Yue, Yidan Liu, Jiayi Ma, Pedram Ghamisi, Antonio Plaza, Leyuan Fang

机构 * School of Artificial Intelligence and Robotics, Hunan University(湖南大学人工智能与机器人学院) Institute of Remote Sensing and Geographic Information System, Peking University(北京大学遥感与地理信息系统研究所) School of Automation, Central South University(中南大学自动化学院) Electronic Information School, Wuhan University(武汉大学电子信息学院) Helmholtz-Zentrum Dresden-Rossendorf(德累斯顿-罗斯托克亥姆霍尔茨中心) Lancaster Environment Centre, Lancaster University(兰卡斯特大学环境研究中心) Hyperspectral Computing Laboratory, Department of Technology of Computers and Communications, Escuela Politécnica, University of Extremadura(埃斯特雷马杜拉大学技术计算机与通讯系超光谱计算实验室)

专题命中 隐私与版权 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 安全评测 25 篇

2510.23334 2025-10-28 cs.CL 83%

Adaptive Blockwise Search: Inference-Time Alignment for Large Language Models

Mohammad Atif Quamar, Mohammad Areeb, Nishant Sharma, Ananth Shreekumar, Jonathan Rosenthal, Muslum Ozgur Ozmen, Mikhail Kuznetsov, Z. Berkay Celik

机构 * Purdue University(普渡大学) Arizona State University(亚利桑那州立大学) Amazon(亚马逊)

专题命中 安全评测 :alignment(title,abstract);harmlessness(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.19674 2025-10-28 cs.CR cs.AI 83%

SAGE: A Generic Framework for LLM Safety Evaluation

Madhur Jindal, Hari Shrawgi, Parag Agrawal, Sandipan Dandapat

专题命中 安全评测 :safety(title,abstract);AI safety(abstract);分类 cs.AI

Comments Accepted to EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22095 2025-10-28 cs.AI cs.CL 81%

Embracing Trustworthy Brain-Agent Collaboration as Paradigm Extension for Intelligent Assistive Technologies

Yankai Chen, Xinni Zhang, Yifei Zhang, Yangning Li, Henry Peng Zou, Chunyu Miao, Weizhi Zhang, Xue Liu, Philip S. Yu

机构 * University of Illinois Chicago(伊利诺伊大学芝加哥分校) MBZUAI(马克斯·普朗克人工智能研究所) McGill University(麦吉尔大学) The Chinese University of Hong Kong(香港中文大学) Nanyang Technological University(南洋理工大学) Tsinghua University(清华大学)

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.CL、cs.AI

Comments Accepted by NeurIPS'25 Position Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21566 2025-10-28 cs.MA cs.CL 79%

ColorEcosystem: Powering Personalized, Standardized, and Trustworthy Agentic Service in massive-agent Ecosystem

Fangwen Wu, Zheng Wu, Jihong Wang, Yunku Chen, Ruiguang Pei, Heyuan Huang, Xin Liao, Xingyu Lou, Huarong Deng, Zhihui Fu, Weiwen Liu, Zhuosheng Zhang, Weinan Zhang, Jun Wang

机构 * Shanghai Jiao Tong University(上海交通大学) OPPO

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.00775 2025-10-28 cs.HC 78%

Efficiency with Rigor! A Trustworthy LLM-powered Workflow for Qualitative Data Analysis

Jie Gao, Zhiyao Shu, Shun Yi Yeo, Alok Prakash, Chien-Ming Huang, Mark Dredze, Ziang Xiao

专题命中 安全评测 :trustworthy(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.06185 2025-10-28 cs.CL cs.AI cs.CY cs.HC cs.LG 70%

Can Large Language Models Unlock Novel Scientific Research Ideas?

Sandeep Kumar, Tirthankar Ghosal, Vinayak Goyal, Asif Ekbal

机构 * Department of Computer Science and Engineering, Indian Institute of Technology Patna(计算机科学与工程系,印度理工学院帕纳布分校) National Center for Computational Sciences, Oak Ridge National Laboratory(计算科学国家中心,橡树岭国家实验室)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI、cs.CY

Comments EMNLP 2025 (Main)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22728 2025-10-28 cs.LG cs.CV 70%

S-Chain: Structured Visual Chain-of-Thought For Medicine

Khai Le-Duc, Duy M. H. Nguyen, Phuong T. H. Trinh, Tien-Phat Nguyen, Nghiem T. Diep, An Ngo, Tung Vu, Trinh Vuong, Anh-Tien Nguyen, Mau Nguyen, Van Trung Hoang, Khai-Nguyen Nguyen, Hy Nguyen, Chris Ngo, Anji Liu, Nhat Ho, Anne-Christin Hauschild, Khanh Xuan Nguyen, Thanh Nguyen-Tang, Pengtao Xie, Daniel Sonntag, James Zou, Mathias Niepert, Anh Totti Nguyen

专题命中 安全评测 :alignment(abstract);trustworthy(abstract);分类 cs.LG

Comments First version

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23471 2025-10-28 stat.ML cs.AI cs.LG 62%

Robust Decision Making with Partially Calibrated Forecasts

Shayan Kiyani, Hamed Hassani, George Pappas, Aaron Roth

机构 * University of Pennsylvania(宾夕法尼亚大学)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.19209 2025-10-28 cs.CL cs.AI cs.CE stat.ML 62%

MOOSE-Chem2: Exploring LLM Limits in Fine-Grained Scientific Hypothesis Discovery via Hierarchical Search

Zonglin Yang, Wanhao Liu, Ben Gao, Yujie Liu, Wei Li, Tong Xie, Lidong Bing, Wanli Ouyang, Erik Cambria, Dongzhan Zhou

机构 * Nanyang Technological University(南洋理工大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) University of Science and Technology of China(中国科学技术大学) Wuhan University(武汉大学) National University of Singapore(新加坡国立大学) University of New South Wales(新南威尔士大学) MiroMind The Chinese University of Hong Kong(香港中文大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

Comments Accepted by NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.17823 2025-10-28 cs.CV cs.AI cs.LG 62%

Robust Multimodal Learning via Cross-Modal Proxy Tokens

Md Kaykobad Reza, Ameya Patil, Mashhour Solh, M. Salman Asif

机构 * University of California Riverside(加州大学河滨分校) Amazon(亚马逊)

专题命中 安全评测 :alignment(abstract);分类 cs.AI、cs.LG

Comments 28 Pages, 13 Figures, 11 Tables. Accepted by Transactions on Machine Learning Research (TMLR)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.09947 2025-10-28 cs.LG cs.AI 62%

Identifying Trustworthiness Challenges in Deep Learning Models for Continental-Scale Water Quality Prediction

Xiaobo Xia, Xiaofeng Liu, Jiale Liu, Kuai Fang, Lu Lu, Samet Oymak, William S. Currie, Tongliang Liu

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG

Comments Accepted by Nexus (Cell Press). 61 pages, 24 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22998 2025-10-28 cs.AI 57%

ProfileXAI: User-Adaptive Explainable AI

Gilber A. Corrales, Carlos Andrés Ferro Sánchez, Reinel Tabares-Soto, Jesús Alfonso López Sotelo, Gonzalo A. Ruz, Johan Sebastian Piña Durán

机构 * Facultad de Ingeniería y Ciencias Básicas, Universidad Autónoma de Occidente(工程与基础科学学院,自治大学)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

Comments pages, 1 figure, 3 tables. Preprint. Evaluated on UCI Heart Disease (1989) and UCI Differentiated Thyroid Cancer Recurrence (2023). Uses IEEEtran

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.16343 2025-10-28 cs.SD cs.AI eess.AS 57%

Detect Any Sound: Open-Vocabulary Sound Event Detection with Multi-Modal Queries

Pengfei Cai, Yan Song, Qing Gu, Nan Jiang, Haoyu Song, Ian McLoughlin

机构 * University of Science and Technology of China(科学技术大学) Singapore Institute of Technology(新加坡理工学院)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

Comments Accepted by MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23035 2025-10-28 cs.CR cs.AI 57%

A high-capacity linguistic steganography based on entropy-driven rank-token mapping

Jun Jiang, Weiming Zhang, Nenghai Yu, Kejiang Chen

机构 * School of Cyber Science and Technology(网络科学与技术学院) University of Science and Technology of China(中国科学技术大学)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22618 2025-10-28 cs.CV cs.AI 57%

Cross-Species Transfer Learning in Agricultural AI: Evaluating ZebraPose Adaptation for Dairy Cattle Pose Estimation

Mackenzie Tapp, Sibi Chakravarthy Parivendan, Kashfia Sailunaz, Suresh Neethirajan

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

Comments 20 pages, 11 figures, 6 Tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.20414 2025-10-28 cs.GR cs.CV cs.LG cs.RO 57%

SceneWeaver: All-in-One 3D Scene Synthesis with an Extensible and Self-Reflective Agent

Yandan Yang, Baoxiong Jia, Shujie Zhang, Siyuan Huang

机构 * State Key Laboratory of General Artificial Intelligence, BIGAI(1 通用人工智能国家重点实验室、BIGAI) Tsinghua University(2 清华大学)

专题命中 安全评测 :alignment(abstract);分类 cs.LG

Comments Accepted by NeurIPS 2025, 26 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.23881 2025-10-28 cs.CV cs.LG 57%

Spurious-Aware Prototype Refinement for Reliable Out-of-Distribution Detection

Reihaneh Zohrabi, Hosein Hasani, Mahdieh Soleymani Baghshah, Anna Rohrbach, Marcus Rohrbach, Mohammad Hossein Rohban

机构 * TU Darmstadt(图宾根大学) Sharif University of Technology(谢赫·穆罕默德·卡齐姆大学)

专题命中 安全评测 :safety(abstract);分类 cs.LG

Comments Accepted at NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.12925 2025-10-28 cs.SE cs.AI cs.IR 57%

CPRet: A Dataset, Benchmark, and Model for Retrieval in Competitive Programming

Han Deng, Yuan Meng, Shixiang Tang, Wanli Ouyang, Xinzhu Ma

机构 * The Chinese University of Hong Kong(香港中文大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Tsinghua University(清华大学) Beihang University(北航) Shenzhen Loop Area Institute(深圳河套学院)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

Comments Accepted by NeurIPS 2025 Dataset and Benchmark Track

详情

展开后加载摘要…

URL PDF HTML 收藏