arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-10-03 至 2025-10-03 共收录 22 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 22 篇

2510.01990 2025-10-03 cs.CV cs.CY 86%

TriAlignXA: An Explainable Trilemma Alignment Framework for Trustworthy Agri-product Grading

Jianfei Xie, Ziyang Li

机构 * School of Software, Xinjiang University(软件学院,新疆大学)

专题命中 安全评测 :trustworthy(title,abstract);alignment(title);分类 cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01606 2025-10-03 cs.IR cs.AI cs.CL 81%

Bridging Collaborative Filtering and Large Language Models with Dynamic Alignment, Multimodal Fusion and Evidence-grounded Explanations

Bo Ma, LuYao Liu, Simon Lau, Chandler Yuan, and XueY Cui, Rosie Zhang

机构 * Department of Software \& Microelectronics, Peking University, Beijing, China Economic Law School, China University of Political Science Financial Media, Peking University, ChangSha, China

专题命中 安全评测 :alignment(title);trustworthy(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01231 2025-10-03 cs.CL cs.AI stat.ML 81%

Trustworthy Summarization via Uncertainty Quantification and Risk Awareness in Large Language Models

Shuaidong Pan, Di Wu

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.00577 2025-10-03 cs.AI 79%

A Flexible Method for Behaviorally Measuring Alignment Between Human and Artificial Intelligence Using Representational Similarity Analysis

Mattson Ogg, Ritwik Bose, Jamie Scharf, Christopher Ratto, Michael Wolmetz

专题命中 安全评测 :alignment(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01780 2025-10-03 cs.CR cs.AI cs.CY cs.LG 75%

Secure Multi-Modal Data Fusion in Federated Digital Health Systems via MCP

Aueaphum Aueawatthanaphisut

机构 * School of Information, Computer, and Communication Technology(信息、计算机与通信技术学院) Sirindhorn International Institute of Technology, Thammasat University(泰国朱拉安吞国际技术学院,泰国 Thammasat 大学)

专题命中 安全评测 :alignment(abstract);trustworthy(abstract);分类 cs.AI、cs.CY、cs.LG

Comments 6 pages, 8 figures, 7 equations, 1 algorithm

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02204 2025-10-03 cs.CL 70%

Say One Thing, Do Another? Diagnosing Reasoning-Execution Gaps in VLM-Powered Mobile-Use Agents

Lingzhong Dong, Ziqi Zhou, Shuaibo Yang, Haiyue Sheng, Pengzhou Cheng, Zongru Wu, Zheng Wu, Gongshen Liu, Zhuosheng Zhang

机构 * Shanghai Jiao Tong University(上海交通大学) Beijing Institute of Technology(北京理工大学)

专题命中 安全评测 :alignment(abstract);trustworthy(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01815 2025-10-03 cs.AI 70%

Human-AI Teaming Co-Learning in Military Operations

Clara Maathuis, Kasper Cools

专题命中 安全评测 :safety(abstract);trustworthy(abstract);分类 cs.AI

Comments Submitted to Sensors + Imaging; presented on 18th of September (Artificial Intelligence for Security and Defence Applications III)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01670 2025-10-03 cs.AI cs.CL cs.CR cs.CY cs.LG 70%

Just Do It!? Computer-Use Agents Exhibit Blind Goal-Directedness

Erfan Shayegani, Keegan Hines, Yue Dong, Nael Abu-Ghazaleh, Roman Lutz, Spencer Whitehead, Vidhisha Balachandran, Besmira Nushi, Vibhav Vineet

机构 * Microsoft Research AI Frontiers(微软研究院人工智能前沿) Microsoft AI Red Team(微软AI红色团队) University of California, Riverside(加州大学河滨分校) NVIDIA(英伟达)

专题命中 安全评测 :safety(abstract);分类 cs.CL、cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01691 2025-10-03 cs.CV 67%

MedQ-Bench: Evaluating and Exploring Medical Image Quality Assessment Abilities in MLLMs

Jiyao Liu, Jinjie Wei, Wanying Qu, Chenglong Ma, Junzhi Ning, Yunheng Li, Ying Chen, Xinzhe Luo, Pengcheng Chen, Xin Gao, Ming Hu, Huihui Xu, Xin Wang, Shujian Gao, Dingkang Yang, Zhongying Deng, Jin Ye, Lihao Liu, Junjun He, Ningsheng Xu

机构 * Fudan University(复旦大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Imperial College London(帝国理工学院) University of Cambridge(剑桥大学)

专题命中 安全评测 :alignment(abstract);safety(abstract)

Comments 26 pages, 13 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01792 2025-10-03 cs.CL cs.AI cs.IR 62%

Comparison of Unsupervised Metrics for Evaluating Judicial Decision Extraction

Ivan Leonidovich Litvak, Anton Kostin, Fedor Lashkin, Tatiana Maksiyan, Sergey Lagutin

机构 * Moscow Center for Advanced Studies(莫斯科高级研究学院)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

Comments 28 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01270 2025-10-03 cs.CL cs.AI 62%

Think Twice, Generate Once: Safeguarding by Progressive Self-Reflection

Hoang Phan, Victor Li, Qi Lei

机构 * New York University(纽约大学)

专题命中 安全评测 :safety(abstract);分类 cs.CL、cs.AI

Comments Accepted to EMNLP 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25264 2025-10-03 cs.DB cs.AI cs.LG cs.SE 62%

GeoSQL-Eval: First Evaluation of LLMs on PostGIS-Based NL2GeoSQL Queries

Shuyang Hou, Haoyue Jiao, Ziqi Liu, Lutong Xie, Guanyu Chen, Shaowen Wu, Xuefeng Guan, Huayi Wu

专题命中 安全评测 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01671 2025-10-03 cs.AI cs.HC 57%

A Locally Executable AI System for Improving Preoperative Patient Communication: A Multi-Domain Clinical Evaluation

Motoki Sato, Yuki Matsushita, Hidekazu Takahashi, Tomoaki Kakazu, Sou Nagata, Mizuho Ohnuma, Atsushi Yoshikawa, Masayuki Yamamura

专题命中 安全评测 :safety(abstract);分类 cs.AI

Comments 32 pages, 4 figures, 10 tables 32 pages, 4 figures, 10 tables. This paper is currently under review at ACM Transactions on Computing for Healthcare. Reproducibility resources: http://github.com/motokinaru/LENOHA-medical-dialogue

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01295 2025-10-03 cs.AI cs.MA 57%

The Social Laboratory: A Psychometric Framework for Multi-Agent LLM Evaluation

Zarreen Reza

机构 * Independent researcher(独立研究者)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

Comments 39th Conference on Neural Information Processing Systems (NeurIPS 2025) Workshop on Evaluating the Evolving LLM Lifecycle: Benchmarks, Emergent Abilities, and Scaling

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.26301 2025-10-03 cs.LG cs.HC 57%

NeuroTTT: Bridging Pretraining-Downstream Task Misalignment in EEG Foundation Models via Test-Time Training

Suli Wang, Yangshen Deng, Zhenghua Bao, Xinyu Zhan, Yiqun Duan

机构 * Technical University of Darmstadt(达姆斯塔特技术大学) University of Edinburgh(爱丁堡大学) University of Technology Sydney(悉尼技术大学)

专题命中 安全评测 :alignment(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.21285 2025-10-03 cs.CL 57%

Double-Checker: Enhancing Reasoning of Slow-Thinking LLMs via Self-Critical Fine-Tuning

Xin Xu, Tianhao Chen, Fan Zhang, Wanlong Liu, Pengxiang Li, Ajay Kumar Jaiswal, Yuchen Yan, Jishan Hu, Yang Wang, Hao Chen, Shiwei Liu, Shizhe Diao, Can Yang, Lu Yin

机构 * The Hong Kong University of Science and Technology(香港科学与技术大学) University of Electronic Science and Technology of China(电子科技大学) Dalian University of Technology(大连理工大学) University of Texas at Austin(德克萨斯大学奥斯汀分校) Zhejiang University(浙江大学) University of Oxford(牛津大学) NVIDIA(NVIDIA公司) University of Surrey(萨里大学)

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL

Comments 10 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.02064 2025-10-03 cs.CY cs.HC 57%

The Measurement Imbalance in Agentic AI Evaluation Undermines Industry Productivity Claims

Kiana Jafari Meimandi, Gabriela Aránguiz-Dias, Grace Ra Kim, Lana Saadeddin, Allie Griffith, Mykel J. Kochenderfer

专题命中 安全评测 :safety(abstract);分类 cs.CY

Comments 15 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.08198 2025-10-03 stat.ML cs.LG 57%

SIM-Shapley: A Stable and Computationally Efficient Approach to Shapley Value Approximation

Wangxuan Fan, Siqi Li, Doudou Zhou, Yohei Okada, Chuan Hong, Molei Liu, Nan Liu

机构 * National University of Singapore(新加坡国立大学) Duke-NUS Medical School(杜克-国立新加坡大学医学院) Duke University(杜克大学) Peking University(北京大学)

专题命中 安全评测 :trustworthy(abstract);分类 cs.LG

Comments 21 pages, 6 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.12051 2025-10-03 cs.CL 57%

TLUE: A Tibetan Language Understanding Evaluation Benchmark

Fan Gao, Cheng Huang, Nyima Tashi, Xiangxiang Wang, Thupten Tsering, Ban Ma-bao, Renzeg Duojie, Gadeng Luosang, Rinchen Dongrub, Dorje Tashi, Hao Wang Xiao Feng, Yongbin Yu

机构 * University of Electronic Science and Technology of China(电子科技大学) Tibet University(西藏大学) University of Texas Southwestern Medical Center(德克萨斯西南医学中心) Southern Methodist University(南方 Methodist 大学) The State Key Laboratory of Tibetan Intelligence(藏语智能国家重点实验室) University of Connecticut(康涅狄格大学)

专题命中 安全评测 :safety(abstract);分类 cs.CL

Comments Accepted by EMNLP Main Conference (Poster)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02169 2025-10-03 cs.SE cs.CR 50%

TAIBOM: Bringing Trustworthiness to AI-Enabled Systems

Vadim Safronov, Anthony McCaigue, Nicholas Allott, Andrew Martin

专题命中 安全评测 :trustworthy(abstract)

Comments This paper has been accepted at the First International Workshop on Security and Privacy-Preserving AI/ML (SPAIML 2025), co-located with the 28th European Conference on Artificial Intelligence (ECAI 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01683 2025-10-03 cs.CV 50%

Uncovering Overconfident Failures in CXR Models via Augmentation-Sensitivity Risk Scoring

Han-Jay Shu, Wei-Ning Chiu, Shun-Ting Chang, Meng-Ping Huang, Takeshi Tohyama, Ahram Han, Po-Chih Kuo

机构 * National Tsing Hua University(国立清华大学) National Taiwan University(国立台湾大学) Massachusetts Institute of Technology(麻省理工学院)

专题命中 安全评测 :safety(abstract)

Comments 5 pages, 1 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04677 2025-10-03 cs.CV 50%

Robust Prompt Tuning for Vision-Language Models with Mild Semantic Noise

Yansheng Gao, Yufei Zheng, Shengsheng Wang

机构 * College of Computer Science and Technology, Key Laboratory of Symbolic Computation and Knowledge Engineering of Ministry of Education, Jilin University(计算机科学与技术学院、教育部符号计算与知识工程重点实验室、吉林大学) College of Software, Key Laboratory of Symbolic Computation and Knowledge Engineering of Ministry of Education, Jilin University(软件学院、教育部符号计算与知识工程重点实验室、吉林大学)

专题命中 安全评测 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏