arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-10-03 至 2025-10-03 共收录 61 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 22 篇

2510.01780 2025-10-03 cs.CR cs.AI cs.CY cs.LG 75%

Secure Multi-Modal Data Fusion in Federated Digital Health Systems via MCP

Aueaphum Aueawatthanaphisut

机构 * School of Information, Computer, and Communication Technology(信息、计算机与通信技术学院) Sirindhorn International Institute of Technology, Thammasat University(泰国朱拉安吞国际技术学院,泰国 Thammasat 大学)

专题命中 安全评测 :alignment(abstract);trustworthy(abstract);分类 cs.AI、cs.CY、cs.LG

Comments 6 pages, 8 figures, 7 equations, 1 algorithm

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02204 2025-10-03 cs.CL 70%

Say One Thing, Do Another? Diagnosing Reasoning-Execution Gaps in VLM-Powered Mobile-Use Agents

Lingzhong Dong, Ziqi Zhou, Shuaibo Yang, Haiyue Sheng, Pengzhou Cheng, Zongru Wu, Zheng Wu, Gongshen Liu, Zhuosheng Zhang

机构 * Shanghai Jiao Tong University(上海交通大学) Beijing Institute of Technology(北京理工大学)

专题命中 安全评测 :alignment(abstract);trustworthy(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01815 2025-10-03 cs.AI 70%

Human-AI Teaming Co-Learning in Military Operations

Clara Maathuis, Kasper Cools

专题命中 安全评测 :safety(abstract);trustworthy(abstract);分类 cs.AI

Comments Submitted to Sensors + Imaging; presented on 18th of September (Artificial Intelligence for Security and Defence Applications III)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01670 2025-10-03 cs.AI cs.CL cs.CR cs.CY cs.LG 70%

Just Do It!? Computer-Use Agents Exhibit Blind Goal-Directedness

Erfan Shayegani, Keegan Hines, Yue Dong, Nael Abu-Ghazaleh, Roman Lutz, Spencer Whitehead, Vidhisha Balachandran, Besmira Nushi, Vibhav Vineet

机构 * Microsoft Research AI Frontiers(微软研究院人工智能前沿) Microsoft AI Red Team(微软AI红色团队) University of California, Riverside(加州大学河滨分校) NVIDIA(英伟达)

专题命中 安全评测 :safety(abstract);分类 cs.CL、cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01691 2025-10-03 cs.CV 67%

MedQ-Bench: Evaluating and Exploring Medical Image Quality Assessment Abilities in MLLMs

Jiyao Liu, Jinjie Wei, Wanying Qu, Chenglong Ma, Junzhi Ning, Yunheng Li, Ying Chen, Xinzhe Luo, Pengcheng Chen, Xin Gao, Ming Hu, Huihui Xu, Xin Wang, Shujian Gao, Dingkang Yang, Zhongying Deng, Jin Ye, Lihao Liu, Junjun He, Ningsheng Xu

机构 * Fudan University(复旦大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Imperial College London(帝国理工学院) University of Cambridge(剑桥大学)

专题命中 安全评测 :alignment(abstract);safety(abstract)

Comments 26 pages, 13 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01792 2025-10-03 cs.CL cs.AI cs.IR 62%

Comparison of Unsupervised Metrics for Evaluating Judicial Decision Extraction

Ivan Leonidovich Litvak, Anton Kostin, Fedor Lashkin, Tatiana Maksiyan, Sergey Lagutin

机构 * Moscow Center for Advanced Studies(莫斯科高级研究学院)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

Comments 28 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01270 2025-10-03 cs.CL cs.AI 62%

Think Twice, Generate Once: Safeguarding by Progressive Self-Reflection

Hoang Phan, Victor Li, Qi Lei

机构 * New York University(纽约大学)

专题命中 安全评测 :safety(abstract);分类 cs.CL、cs.AI

Comments Accepted to EMNLP 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25264 2025-10-03 cs.DB cs.AI cs.LG cs.SE 62%

GeoSQL-Eval: First Evaluation of LLMs on PostGIS-Based NL2GeoSQL Queries

Shuyang Hou, Haoyue Jiao, Ziqi Liu, Lutong Xie, Guanyu Chen, Shaowen Wu, Xuefeng Guan, Huayi Wu

专题命中 安全评测 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01671 2025-10-03 cs.AI cs.HC 57%

A Locally Executable AI System for Improving Preoperative Patient Communication: A Multi-Domain Clinical Evaluation

Motoki Sato, Yuki Matsushita, Hidekazu Takahashi, Tomoaki Kakazu, Sou Nagata, Mizuho Ohnuma, Atsushi Yoshikawa, Masayuki Yamamura

专题命中 安全评测 :safety(abstract);分类 cs.AI

Comments 32 pages, 4 figures, 10 tables 32 pages, 4 figures, 10 tables. This paper is currently under review at ACM Transactions on Computing for Healthcare. Reproducibility resources: http://github.com/motokinaru/LENOHA-medical-dialogue

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01295 2025-10-03 cs.AI cs.MA 57%

The Social Laboratory: A Psychometric Framework for Multi-Agent LLM Evaluation

Zarreen Reza

机构 * Independent researcher(独立研究者)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

Comments 39th Conference on Neural Information Processing Systems (NeurIPS 2025) Workshop on Evaluating the Evolving LLM Lifecycle: Benchmarks, Emergent Abilities, and Scaling

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.26301 2025-10-03 cs.LG cs.HC 57%

NeuroTTT: Bridging Pretraining-Downstream Task Misalignment in EEG Foundation Models via Test-Time Training

Suli Wang, Yangshen Deng, Zhenghua Bao, Xinyu Zhan, Yiqun Duan

机构 * Technical University of Darmstadt(达姆斯塔特技术大学) University of Edinburgh(爱丁堡大学) University of Technology Sydney(悉尼技术大学)

专题命中 安全评测 :alignment(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.21285 2025-10-03 cs.CL 57%

Double-Checker: Enhancing Reasoning of Slow-Thinking LLMs via Self-Critical Fine-Tuning

Xin Xu, Tianhao Chen, Fan Zhang, Wanlong Liu, Pengxiang Li, Ajay Kumar Jaiswal, Yuchen Yan, Jishan Hu, Yang Wang, Hao Chen, Shiwei Liu, Shizhe Diao, Can Yang, Lu Yin

机构 * The Hong Kong University of Science and Technology(香港科学与技术大学) University of Electronic Science and Technology of China(电子科技大学) Dalian University of Technology(大连理工大学) University of Texas at Austin(德克萨斯大学奥斯汀分校) Zhejiang University(浙江大学) University of Oxford(牛津大学) NVIDIA(NVIDIA公司) University of Surrey(萨里大学)

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL

Comments 10 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.02064 2025-10-03 cs.CY cs.HC 57%

The Measurement Imbalance in Agentic AI Evaluation Undermines Industry Productivity Claims

Kiana Jafari Meimandi, Gabriela Aránguiz-Dias, Grace Ra Kim, Lana Saadeddin, Allie Griffith, Mykel J. Kochenderfer

专题命中 安全评测 :safety(abstract);分类 cs.CY

Comments 15 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.08198 2025-10-03 stat.ML cs.LG 57%

SIM-Shapley: A Stable and Computationally Efficient Approach to Shapley Value Approximation

Wangxuan Fan, Siqi Li, Doudou Zhou, Yohei Okada, Chuan Hong, Molei Liu, Nan Liu

机构 * National University of Singapore(新加坡国立大学) Duke-NUS Medical School(杜克-国立新加坡大学医学院) Duke University(杜克大学) Peking University(北京大学)

专题命中 安全评测 :trustworthy(abstract);分类 cs.LG

Comments 21 pages, 6 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.12051 2025-10-03 cs.CL 57%

TLUE: A Tibetan Language Understanding Evaluation Benchmark

Fan Gao, Cheng Huang, Nyima Tashi, Xiangxiang Wang, Thupten Tsering, Ban Ma-bao, Renzeg Duojie, Gadeng Luosang, Rinchen Dongrub, Dorje Tashi, Hao Wang Xiao Feng, Yongbin Yu

机构 * University of Electronic Science and Technology of China(电子科技大学) Tibet University(西藏大学) University of Texas Southwestern Medical Center(德克萨斯西南医学中心) Southern Methodist University(南方 Methodist 大学) The State Key Laboratory of Tibetan Intelligence(藏语智能国家重点实验室) University of Connecticut(康涅狄格大学)

专题命中 安全评测 :safety(abstract);分类 cs.CL

Comments Accepted by EMNLP Main Conference (Poster)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02169 2025-10-03 cs.SE cs.CR 50%

TAIBOM: Bringing Trustworthiness to AI-Enabled Systems

Vadim Safronov, Anthony McCaigue, Nicholas Allott, Andrew Martin

专题命中 安全评测 :trustworthy(abstract)

Comments This paper has been accepted at the First International Workshop on Security and Privacy-Preserving AI/ML (SPAIML 2025), co-located with the 28th European Conference on Artificial Intelligence (ECAI 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01683 2025-10-03 cs.CV 50%

Uncovering Overconfident Failures in CXR Models via Augmentation-Sensitivity Risk Scoring

Han-Jay Shu, Wei-Ning Chiu, Shun-Ting Chang, Meng-Ping Huang, Takeshi Tohyama, Ahram Han, Po-Chih Kuo

机构 * National Tsing Hua University(国立清华大学) National Taiwan University(国立台湾大学) Massachusetts Institute of Technology(麻省理工学院)

专题命中 安全评测 :safety(abstract)

Comments 5 pages, 1 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04677 2025-10-03 cs.CV 50%

Robust Prompt Tuning for Vision-Language Models with Mild Semantic Noise

Yansheng Gao, Yufei Zheng, Shengsheng Wang

机构 * College of Computer Science and Technology, Key Laboratory of Symbolic Computation and Knowledge Engineering of Ministry of Education, Jilin University(计算机科学与技术学院、教育部符号计算与知识工程重点实验室、吉林大学) College of Software, Key Laboratory of Symbolic Computation and Knowledge Engineering of Ministry of Education, Jilin University(软件学院、教育部符号计算与知识工程重点实验室、吉林大学)

专题命中 安全评测 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

2. AI治理与伦理 2 篇

2509.22872 2025-10-03 cs.CY 89%

Anti-Regulatory AI: How "AI Safety" is Leveraged Against Regulatory Oversight

Rui-Jie Yew, Brian Judge

专题命中 AI治理与伦理 :safety(title,abstract);AI safety(title,abstract);alignment(abstract);分类 cs.CY

Comments Forthcoming at EAAMO 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01963 2025-10-03 cs.SD cs.LG 57%

Bias beyond Borders: Global Inequalities in AI-Generated Music

Ahmet Solak, Florian Grötschla, Luca A. Lanzendörfer, Roger Wattenhofer

机构 * ETH Zurich(苏黎世联邦理工学院)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 其他安全 11 篇

2510.01924 2025-10-03 cs.AI cs.MA 79%

To Mask or to Mirror: Human-AI Alignment in Collective Reasoning

Crystal Qian, Aaron Parisi, Clémentine Bouleau, Vivian Tsai, Maël Lebreton, Lucas Dixon

机构 * Google DeepMind(谷歌DeepMind) Paris School of Economics(巴黎经济学院)

专题命中 其他安全 :alignment(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01454 2025-10-03 cs.CV cs.LG 74%

Data Selection for Fine-tuning Vision Language Models via Cross Modal Alignment Trajectories

Nilay Naharas, Dang Nguyen, Nesihan Bulut, Mohammadhossein Bateni, Vahab Mirrokni, Baharan Mirzasoleiman

机构 * Department of Computer Science, University of California Los Angeles(加州大学洛杉矶分校计算机科学系) Google Research(谷歌研究)

专题命中 其他安全 :alignment(title);分类 cs.LG

Comments 30 pages, 10 figures, 5 tables, link: https://bigml-cs-ucla.github.io/XMAS-project-page/

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01428 2025-10-03 q-bio.QM cs.AI 74%

BioVERSE: Representation Alignment of Biomedical Modalities to LLMs for Multi-Modal Reasoning

Ching-Huei Tsou, Michal Ozery-Flato, Ella Barkan, Diwakar Mahajan, Ben Shapira

机构 * IBM T.J. Watson Research Center(IBM T.J. Watson研究所以) IBM Research(IBM研究所以)

专题命中 其他安全 :alignment(title);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01363 2025-10-03 cs.AI 70%

Retrieval-Augmented Framework for LLM-Based Clinical Decision Support

Leon Garza, Anantaa Kotal, Michael A. Grasso, Emre Umucu

机构 * Dept. of Computer Science, The University of Texas at El Paso, USA(计算机科学系,德克萨斯大学埃尔帕索分校) Dept. of Emergency Medicine, University of Maryland School of Medicine, USA(急诊医学系,马里兰大学医学院) Dept. of Public Health Sciences, The University of Texas at El Paso, USA(公共卫生科学系,德克萨斯大学埃尔帕索分校)

专题命中 其他安全 :alignment(abstract);safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18216 2025-10-03 cs.AI cs.LG 62%

nDNA -- the Semantic Helix of Artificial Cognition

Amitava Das

机构 * BITS Pilani, Goa, India(印度戈阿学院)

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00631 2025-10-03 cs.LG cs.AI physics.ao-ph 62%

Forecasting the Ionosphere from Sparse GNSS Data with Temporal-Fusion Transformers

Giacomo Acciarini, Simone Mestici, Halil Kelebek, Linnea Wolniewicz, Michael Vergalla, Madhulika Guhathakurta, Umaa Rebbapragada, Bala Poduval, Atılım Güneş Baydin, Frank Soboczenski

机构 * Advanced Concepts Team European Space Agency(欧洲航天局高级概念团队) Department of Physics Università degli Studi di Roma Sapienza(罗马大学物理系) Department of Engineering Science University of Oxford(牛津大学工程科学系) Department of Information and Computer Science University of Hawai’i at Mānoa(夏威夷大学信息与计算机科学系) Free Flight Research Lab(自由飞行研究实验室) NASA Headquarters(美国国家航空航天局总部) NASA Jet Propulsion Laboratory(美国国家航空航天局喷气推进实验室) University of New Hampshire(新罕布什尔大学) Department of Computer Science University of Oxford, UK(牛津大学计算机科学系) Department of Computer Science University of York & King’s College London(约克大学计算机科学系及伦敦国王学院计算机科学系)

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02270 2025-10-03 cs.CV cs.AI 57%

microCLIP: Unsupervised CLIP Adaptation via Coarse-Fine Token Fusion for Fine-Grained Image Classification

Sathira Silva, Eman Ali, Chetan Arora, Muhammad Haris Khan

机构 * Mohamed Bin Zayed University of Artificial Intelligence(莫莫德·本·扎耶德人工智能大学) Alexandria University(亚历山大大学) IIT Delhi(德里理工学院)

专题命中 其他安全 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02155 2025-10-03 cs.CV cs.AI 57%

Unlocking Vision-Language Models for Video Anomaly Detection via Fine-Grained Prompting

Shu Zou, Xinyu Tian, Lukas Wesemann, Fabian Waschkowski, Zhaoyuan Yang, Jing Zhang

机构 * Australian National University(澳大利亚国立大学) Maincode GE Research(通用电气研究)

专题命中 其他安全 :safety(abstract);分类 cs.AI

Comments 14 pages, video anomaly detection

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01546 2025-10-03 cs.CV cs.LG 57%

Growing Visual Generative Capacity for Pre-Trained MLLMs

Hanyu Wang, Jiaming Han, Ziyan Yang, Qi Zhao, Shanchuan Lin, Xiangyu Yue, Abhinav Shrivastava, Zhenheng Yang, Hao Chen

机构 * University of Maryland, College Park(马里兰大学学院公园分校) CUHK MMLab(香港大学MMLab) ByteDance(字节跳动)

专题命中 其他安全 :alignment(abstract);分类 cs.LG

Comments Project page: https://hywang66.github.io/bridge/

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.18262 2025-10-03 cs.IR cs.LG 57%

Break the ID-Language Barrier: An Adaption Framework for LLM-based Sequential Recommendation

Xiaohan Yu, Li Zhang, Xin Zhao, Yue Wang

机构 * Huawei, Beijing, China(华为,北京,中国) Institute of Finance Technology, UCL, United Kingdom(金融技术研究所,UCL,英国) Civil, Environmental and Geomatic Engineering, UCL, United Kingdom(土木、环境与地理工程,UCL,英国)

专题命中 其他安全 :alignment(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏