arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 9346 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 9346 篇

2512.11362 2025-12-22 cs.RO 67%

An Anatomy of Vision-Language-Action Models: From Modules to Milestones and Challenges

视觉-语言-动作模型的解剖:从模块到里程碑与挑战

Chao Xu, Suyu Zhang, Yang Liu, Baigui Sun, Weihong Chen, Bo Xu, Qi Liu, Juncheng Wang, Shujun Wang, Shan Luo, Jan Peters, Athanasios V. Vasilakos, Stefanos Zafeiriou, Jiankang Deng

机构 * IROOTECH TECHNOLOGY(IROOTECH技术公司) Wolf 1069 b Lab, Sany Group(Sany集团沃尔夫1069b实验室) Department of Engineering, King’s College London(伦敦国王学院工程系) Hong Kong Polytechnic University(香港理工大学) Computer Science Department of the Technische Universität Darmstadt(德累斯顿技术大学计算机科学系) Department of ICT and Center for AI Research, University of Agder (UiA)(阿格德大学信息与通信技术系及人工智能研究中心) Department of Computing, Imperial College London(伦敦帝国理工学院计算系)

专题命中 安全评测 :safety(abstract);trustworthy(abstract)

AI总结 本文系统分析了视觉-语言-动作模型的核心挑战,从模块构建到里程碑发展,为研究者提供结构化指南和未来研究方向。

Comments project page: https://suyuz1.github.io/VLA-Survey-Anatomy/

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.21112 2025-12-17 cs.AI cs.CE cs.CL cs.CY cs.IR 67%

Optimizing Large Language Models for ESG Activity Detection in Financial Texts

优化大型语言模型以检测金融文本中的ESG活动

Mattia Birti, Andrea Maurino, Francesco Osborne

机构 * Department of Informatics, Systems and Communication, University of Milano-Bicocca(信息学、系统与通信系,米兰-比科卡大学) University of Milano-Bicocca(米兰-比科卡大学) The Open University(开放大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI、cs.CY

AI总结 本文提出通过微调优化大型语言模型,提升金融文本中ESG活动检测的准确性。

Comments Published in the Proceedings of the ACM International Conference on AI in Finance (ICAIF). ACM version

Journal ref Proceedings of the ACM International Conference on AI in Finance (ICAIF), 2024, ACM

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.11412 2025-12-15 cs.CE cs.AI cs.CL cs.LG q-bio.BM 67%

Task-Specific Sparse Feature Masks for Molecular Toxicity Prediction with Chemical Language Models

针对分子毒性预测的特定任务稀疏特征掩码与化学语言模型

Kwun Sy Lee, Jiawei Chen, Fuk Sheng Ford Chung, Tianyu Zhao, Zhenyuan Chen, Debby D. Wang

机构 * School of Science and Technology(科学与技术学院) Hong Kong Metropolitan University(香港 Metropolitan 大学) Faculty of Engineering(工程学院) Hong Kong Polytechnic University(香港理工大学)

专题命中 安全评测 :safety(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本文提出了一种多任务学习框架,通过稀疏注意力机制提升化学分子毒性预测的准确性和可解释性。

Comments 6 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18234 2025-12-15 cs.AI cs.CL cs.LG 67%

The Illusion of Readiness in Health AI

健康AI中的“准备”幻觉

Yu Gu, Jingjing Fu, Xiaodong Liu, Jeya Maria Jose Valanarasu, Noel CF Codella, Reuben Tan, Qianchu Liu, Ying Jin, Sheng Zhang, Jinyu Wang, Rui Wang, Lei Song, Guanghui Qin, Naoto Usuyama, Cliff Wong, Hao Cheng, HoHin Lee, Praneeth Sanapathi, Sarah Hilado, Tristan Naumann, Javier Alvarez-Valle, Jiang Bian, Mu Wei, Khalil Malik, Lidong Zhou, Jianfeng Gao, Eric Horvitz, Matthew P. Lungren, Doug Burger, Eric Topol, Hoifung Poon, Paul Vozila

机构 * Scripps Research Translational Institute(斯克里普斯研究转化研究所)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本文通过对抗性压力测试揭示了健康AI在现实应用中的鲁棒性和推理能力不足,强调了对AI系统责任和真实医疗需求的重视。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.08451 2025-12-10 cs.LG cs.AI cs.CY cs.ET 67%

Biothreat Benchmark Generation Framework for Evaluating Frontier AI Models II: Benchmark Generation Process

生物威胁基准生成框架用于评估前沿AI模型II:基准生成过程

Gary Ackerman, Zachary Kallenborn, Anna Wetzel, Hayley Peterson, Jenna LaTourette, Olivia Shoemaker, Brandon Behlendorf, Sheriff Almakki, Doug Clifford, Noah Sheinbaum

专题命中 安全评测 :red teaming(abstract);分类 cs.AI、cs.CY、cs.LG

AI总结 本文提出生物威胁基准生成框架,通过三种方法生成细菌生物威胁基准数据集,用于评估前沿AI模型的生物安全风险。

Comments 18 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.06751 2025-12-09 cs.CL cs.AI cs.LG 67%

Becoming Experienced Judges: Selective Test-Time Learning for Evaluators

成为经验丰富的法官:用于评估者的选择性测试时间学习

Seungyeon Jwa, Daechul Ahn, Reokyoung Kim, Dongyeop Kang, Jonghyun Choi

机构 * Seoul National University(首尔国立大学) University of Minnesota(明尼苏达大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本文提出选择性测试时间学习框架,使评估者在推理过程中逐步改进,通过自适应元提示提升评估效果,实验证明其在多个基准上的优越性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.16543 2025-12-09 cs.IR cs.AI cs.CL cs.LG 67%

The Oracle and The Prism: A Decoupled and Efficient Framework for Generative Recommendation Explanation

oracle 与 prism:一种解耦且高效的生成推荐解释框架

Jiaheng Zhang, Daqiang Zhang

机构 * Sun Yat-sen University(中山大学) School of Software Engineer Tong ji University(同济大学软件工程学院)

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 prism 通过解耦推荐过程和蒸馏技术,提升可解释推荐系统的效率与可信度。

Comments 12pages,3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.05066 2025-12-05 cs.LG cs.AI cs.CL 67%

Multi-LLM Collaboration for Medication Recommendation

多LLM协作用于药物推荐

Huascar Sanchez, Briland Hitaj, Jules Bergmann, Linda Briesemeister

机构 * Computer Science Laboratory, SRI International(SRI国际计算机科学实验室) University of Maryland St. Joseph Medical Center(马里兰大学圣约瑟夫医疗中心)

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本文提出基于LLM化学的多模型协作方法,通过增强互补性、稳定性和校准性,提高药物推荐的可靠性与可信度。

Comments 8 pages, 5 figures, 1 table

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.04871 2025-12-05 cs.AI cs.CL cs.LG 67%

STELLA: Guiding Large Language Models for Time Series Forecasting with Semantic Abstractions

STELLA: 通过语义抽象引导大型语言模型进行时间序列预测

Junjie Fan, Hongye Zhao, Linduo Wei, Jiayu Rao, Guijia Li, Jiaxin Yuan, Wenqi Xu, Yong Qi

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 STELLA通过动态语义抽象机制,提升大型语言模型在时间序列预测中的表现,实现更精准的长期和短期预测。

Comments This work has been submitted to the IEEE for possible publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.04107 2025-12-05 cs.CY cs.AI cs.HC cs.LG 67%

Rethinking AI Evaluation in Education: The TEACH-AI Framework and Benchmark for Generative AI Assistants

重新思考教育中的AI评估:TEACH-AI框架与生成AI助手的基准测试

Shi Ding, Brian Magerko

机构 * Expressive Machinery Lab(表达性机械实验室) Georgia Institute of Technology(佐治亚理工学院)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.CY、cs.LG

AI总结 本文提出TEACH-AI框架,旨在通过多视角重新定义教育中AI的有效性评估,促进包容性和长期影响。

Comments 6 pages, NeurIPS 2025 Responsible Foundation Models Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.22061 2025-12-01 cs.GT 67%

Aligning with Human Values to Enhance Interaction: An eHMI-Mediated Lane-Changing Negotiation Strategy Using Bayesian Inference

与人类价值观对齐以增强交互:一种使用贝叶斯推理的eHMI介导的变道协商策略

Boyao Peng, Linkun Liu

专题命中 安全评测 :alignment(abstract);safety(abstract)

AI总结 本文提出了一种使用贝叶斯推理的eHMI介导变道协商策略,通过善意欺骗提升交互效率和安全性,但同时也揭示了信任崩溃的伦理风险。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.20733 2025-11-27 cs.CY cs.AI cs.CL cs.HC 67%

InvisibleBench: A Deployment Gate for Caregiving Relationship AI

InvisibleBench:护理关系AI的部署门

Ali Madad

机构 * GiveCare

专题命中 安全评测 :safety(abstract);分类 cs.CL、cs.AI、cs.CY

AI总结 InvisibleBench评估护理关系AI的部署安全性,通过多维度测试揭示模型在危机检测、合规性及创伤导向设计上的差异,强调生产系统中确定性危机路由的重要性。

Comments 29 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.20686 2025-11-27 cs.AI cs.CY cs.LG 67%

AssurAI: Experience with Constructing Korean Socio-cultural Datasets to Discover Potential Risks of Generative AI

AssurAI: 韩国社会文化数据集构建经验以发现生成式AI的潜在风险

Chae-Gyun Lim, Seung-Ho Han, EunYoung Byun, Jeongyun Han, Soohyun Cho, Eojin Joo, Heehyeon Kim, Sieun Kim, Juhoon Lee, Hyunsoo Lee, Dongkun Lee, Jonghwan Hyeon, Yechan Hwang, Young-Jun Lee, Kyeongryul Lee, Minhyeong An, Hyunjun Ahn, Jeongwoo Son, Junho Park, Donggyu Yoon, Taehyung Kim, Jeemin Kim, Dasom Choi, Kwangyoung Lee, Hyunseung Lim, Yeohyun Jung, Jongok Hong, Sooyohn Nam, Joonyoung Park, Sungmin Na, Yubin Choi, Jeanne Choi, Yoojin Hong, Sueun Jang, Youngseok Seo, Somin Park, Seoungung Jo, Wonhye Chae, Yeeun Jo, Eunyoung Kim, Joyce Jiyoung Whang, HwaJung Hong, Joseph Seering, Uichin Lee, Juho Kim, Sunna Choi, Seokyeon Ko, Taeho Kim, Kyunghoon Kim, Myungsik Ha, So Jung Lee, Jemin Hwang, JoonHo Kwak, Ho-Jin Choi

专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.CY、cs.LG

AI总结 AssurAI通过构建高质量的韩语多模态数据集,评估生成式AI的安全性,发现潜在风险,促进更安全的AI发展。

Comments 16 pages, HuggingFace: https://huggingface.co/datasets/TTA01/AssurAI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.12336 2025-11-27 cs.CV 67%

Benchmarking the Trustworthiness in Multimodal LLMs for Video Understanding

对多模态大语言模型在视频理解中的可信度进行基准测试

Youze Wang, Zijun Chen, Ruoyu Chen, Shishen Gu, Wenbo Hu, Jiayang Liu, Yinpeng Dong, Hang Su, Jun Zhu, Meng Wang, Richang Hong

机构 * Hefei University of Technology(合肥工业大学) Tsinghua University(清华大学) Institute of Science Tokyo(东京科学研究所)

专题命中 安全评测 :alignment(abstract);safety(abstract)

AI总结 本研究提出Trust-videoLLMs基准,评估23种视频LLMs在真实性、鲁棒性、安全性和隐私等方面的表现,揭示其在动态场景理解及现实风险缓解中的局限性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.19477 2025-11-26 cs.SE 67%

Building Browser Agents: Architecture, Security, and Practical Solutions

构建浏览器代理:架构、安全与实用解决方案

Aram Vardanyan

专题命中 安全评测 :safety(abstract);prompt injection(abstract)

AI总结 本文提出通过专用工具和编程约束构建安全浏览器代理,实现85%的成功率。

Comments 30 pages, 22 figures. Production architecture and benchmark evaluation of browser agents

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.16410 2025-11-21 cs.SE 67%

Data Annotation Quality Problems in AI-Enabled Perception System Development

人工智能感知系统开发中的数据标注质量问题

Hina Saeeda, Tommy Johansson, Mazen Mohamad, Eric Knauss

专题命中 安全评测 :safety(abstract);trustworthy(abstract)

AI总结 本研究通过多组织案例研究,提出了一种涵盖完整性、准确性和一致性的18种标注错误类型分类法,为人工智能感知系统开发中的数据标注质量提供了共享词汇、诊断工具和可操作指导。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.23982 2025-11-19 cs.CV cs.RO 67%

StyleDrive: Towards Driving-Style Aware Benchmarking of End-To-End Autonomous Driving

Ruiyang Hao, Bowen Jing, Haibao Yu, Zaiqing Nie

机构 * AIR, Tsinghua University(空气动力学研究所,清华大学) King’s College London(伦敦国王学院) The University of Manchester(曼彻斯特大学) The University of Hong Kong(香港大学)

专题命中 安全评测 :alignment(abstract);safety(abstract)

Comments 25 pages, 7 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23229 2025-11-19 cs.CL cs.AI cs.CY 67%

MCTSr-Zero: Self-Reflective Psychological Counseling Dialogues Generation via Principles and Adaptive Exploration

Hao Lu, Yanchi Gu, Haoyuan Huang, Yulin Zhou, Ningxin Zhu, Chen Li

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI、cs.CY

Comments 48 pages, 3 figures. Accepted in AAAI-2026 (Main Technical Track). For code and model, see this https://github.com/JianChengXingYun/Mctsr-Zero

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.15621 2025-11-18 cs.SE 67%

DSCodeBench: A Realistic Benchmark for Data Science Code Generation

Shuyin Ouyang, Dong Huang, Jingwen Guo, Zeyu Sun, Qihao Zhu, Jie M. Zhang

专题命中 安全评测 :alignment(abstract);trustworthy(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11635 2025-11-18 cs.CY cs.AI cs.CL 67%

EduAgentQG: A Multi-Agent Workflow Framework for Personalized Question Generation

Rui Jia, Min Zhang, Fengrui Liu, Bo Jiang, Kun Kuang, Zhongxiang Dai

机构 * East China Normal University(东华大学) Shanghai Institute of Al for Education(上海人工智能教育研究院) The Chinese University of Hong Kong, Shenzhen School of Data Science(香港中文大学(深圳)数据科学学院)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.08824 2025-11-18 cs.RO cs.AI cs.CL cs.CY 67%

LLM-Driven Robots Risk Enacting Discrimination, Violence, and Unlawful Actions

Andrew Hundt, Rumaisa Azeem, Masoumeh Mansouri, Martim Brandão

机构 * Carnegie Mellon University(卡内基梅隆大学) King’s College London(伦敦国王学院) University of Birmingham(伯明翰大学)

专题命中 安全评测 :safety(abstract);分类 cs.CL、cs.AI、cs.CY

Comments Published in International Journal of Social Robotics (2025). 49 pages (65 with references and appendix), 27 Figures, 8 Tables. Andrew Hundt and Rumaisa Azeem are equal contribution co-first authors. The positions of the two co-first authors were swapped from arxiv version 1 with the written consent of all four authors. The Version of Record is available via DOI: 10.1007/s12369-025-01301-x

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10665 2025-11-17 cs.CL cs.AI cs.LG 67%

Guarding the Meaning: Self-Supervised Training for Semantic Robustness in Guard Models

Cristina Pinneri, Christos Louizos

机构 * Qualcomm AI Research(高通人工智能研究)

专题命中 安全评测 :safety(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.08702 2025-11-13 cs.LG cs.AI cs.CR cs.CY 67%

FAIRPLAI: A Human-in-the-Loop Approach to Fair and Private Machine Learning

David Sanchez, Holly Lopez, Michelle Buraczyk, Anantaa Kotal

机构 * Dept. of Computer Science, The University of Texas at El Paso(得克萨斯大学埃尔帕索分校计算机科学系) Dept. of Mathematics, Mountain View High School(山景高中数学系) Dept. of Mathematics, El Paso Independent School District(埃尔帕索独立学区数学系)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.CY、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.04715 2025-11-12 cs.CL cs.AI cs.LG 67%

Selection of LLM Fine-Tuning Data based on Orthogonal Rules

Xiaomin Li, Mingye Gao, Zhiwei Zhang, Chang Yue, Hong Hu

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15144 2025-11-10 cs.AI cs.CL cs.CY 67%

HugAgent: Benchmarking LLMs for Simulation of Individualized Human Reasoning

Chance Jiajie Li, Zhenze Mo, Yuhan Tang, Ao Qu, Jiayi Wu, Kaiya Ivy Zhao, Yulu Gan, Jie Fan, Jiangbo Yu, Hang Jiang, Paul Pu Liang, Jinhua Zhao, Luis Alberto Alonso Pastor, Kent Larson

机构 * MIT Media Lab(MIT媒体实验室) MIT EECS(MIT电子工程与计算机科学系) MIT BCS(MIT生物科学系) MIT IDSS(MIT国际设计系统研究所) MIT CEE(MIT土木与环境工程系) MIT DUSP(MIT设计学教授职位) MIT Architecture(MIT建筑系) Northeastern University(东北大学) Brown University(布朗大学) McGill University(麦吉尔大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI、cs.CY

Comments To appear in NeurIPS 2025 Workshop on Bridging Language, Agent, and World Models (LAW)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.03641 2025-11-06 cs.CR cs.AI cs.CL cs.CY 67%

Watermarking Large Language Models in Europe: Interpreting the AI Act in Light of Technology

Thomas Souverain

机构 * Department of AI Ethics, CEA Paris-Saclay(人工智能伦理系,CEA巴黎萨克雷)

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL、cs.AI、cs.CY

Comments 17 pages, 2 Tables and 2 Pictures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21861 2025-11-06 cs.LG cs.AI cs.CL 67%

The Mirror Loop: Recursive Non-Convergence in Generative Reasoning Systems

Bentley DeVilling

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments 18 pages, 2 figures. Category: cs.LG. Code and data: https://github.com/Course-Correct-Labs/mirror-loop

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.24414 2025-11-05 cs.CV 67%

A Quantitative Evaluation Framework for Explainable AI in Semantic Segmentation

Reem Hammoud, Abdul Karim Gizzini, Ali J. Ghandour

机构 * American University of Beirut(美国贝鲁特美国大学) SogetiLabs Research and Innovation(SogetiLabs研究与创新) National Center for Remote Sensing(远程 sensing 国家中心)

专题命中 安全评测 :safety(abstract);trustworthy(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.24898 2025-10-30 eess.SY cs.SY 67%

Delay Tolerant Control for Autonomous Driving Using CDOB

Xincheng Cao, Haochong Chen, Levent Guvenc, Bilin Aksun-Guvenc

专题命中 安全评测 :alignment(abstract);safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.24811 2025-10-30 cs.CL cs.AI cs.LG 67%

ProofSketch: Efficient Verified Reasoning for Large Language Models

Disha Sheshanarayana, Tanishka Magar

机构 * Manipal University Jaipur(马普尔大学斋普尔)

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL、cs.AI、cs.LG

Comments Accepted at NeurIPS 2025, ER Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏