arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 9380 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 9380 篇

2504.19338 2025-04-29 physics.flu-dyn 71%

OpenFOAMGPT 2.0: end-to-end, trustworthy automation for computational fluid dynamics

Jingsen Feng, Ran Xu, Xu Chu

专题命中 安全评测 :trustworthy(title)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.07189 2025-04-11 eess.SY cs.SY 71%

Multi-Agent Trustworthy Consensus under Random Dynamic Attacks

Orhan Eren Akgün, Sarper Aydın, Stephanie Gil, Angelia Nedić

专题命中 安全评测 :trustworthy(title)

Comments 16 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.19391 2025-03-26 cs.CV cs.MA 71%

TraF-Align: Trajectory-aware Feature Alignment for Asynchronous Multi-agent Perception

Zhiying Song, Lei Yang, Fuxi Wen, Jun Li

专题命中 安全评测 :alignment(title)

Comments Accepted to CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.07906 2025-03-12 cs.CV 71%

Painting with Words: Elevating Detailed Image Captioning with Benchmark and Alignment Learning

Qinghao Ye, Xianhan Zeng, Fu Li, Chunyuan Li, Haoqi Fan

专题命中 安全评测 :alignment(title)

Comments Accepted by ICLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.00663 2025-02-20 cs.CV cs.RO 71%

Generalized Robot 3D Vision-Language Model with Fast Rendering and Pre-Training Vision-Language Alignment

Kangcheng Liu, Yong-Jin Liu, Baoquan Chen

专题命中 安全评测 :alignment(title)

Comments IEEE Transactions on Pattern Analysis and Machine Intelligence, Manuscript Info: 17 Pages, 13 Figures, and 6 Tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.09048 2024-12-17 cs.SE 71%

Towards Trustworthy LLMs for Code: A Data-Centric Synergistic Auditing Framework

Chong Wang, Zhenpeng Chen, Tianlin Li, Yilun Zhao, Yang Liu

专题命中 安全评测 :trustworthy(title)

Comments Short Vision Paper, Accepted by ICSE'25-NIER

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.10122 2024-10-02 cs.CV 71%

Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Bin Lin, Yang Ye, Bin Zhu, Jiaxi Cui, Munan Ning, Peng Jin, Li Yuan

专题命中 安全评测 :alignment(title)

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.16455 2024-09-26 cs.RO 71%

MultiTalk: Introspective and Extrospective Dialogue for Human-Environment-LLM Alignment

Venkata Naren Devarakonda, Ali Umut Kaypak, Shuaihang Yuan, Prashanth Krishnamurthy, Yi Fang, Farshad Khorrami

专题命中 安全评测 :alignment(title)

Comments 7 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.10823 2024-08-21 cs.CV eess.IV 71%

Trustworthy Compression? Impact of AI-based Codecs on Biometrics for Law Enforcement

Sandra Bergmann, Denise Moussa, Christian Riess

专题命中 安全评测 :trustworthy(title)

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.00098 2024-07-02 eess.IV cs.CV 71%

Scalable, Trustworthy Generative Model for Virtual Multi-Staining from H&E Whole Slide Images

Mehdi Ounissi, Ilias Sarbout, Jean-Pierre Hugot, Christine Martinez-Vinson, Dominique Berrebi, Daniel Racoceanu

专题命中 安全评测 :trustworthy(title)

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.01459 2024-01-12 cs.CV 71%

Align Your Prompts: Test-Time Prompting with Distribution Alignment for Zero-Shot Generalization

Jameel Hassan, Hanan Gani, Noor Hussein, Muhammad Uzair Khattak, Muzammal Naseer, Fahad Shahbaz Khan, Salman Khan

专题命中 安全评测 :alignment(title)

Comments Accepted to NeurIPS 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.08533 2024-01-08 cs.CR 71%

Trustchain -- Trustworthy Decentralised Public Key Infrastructure for Digital Credentials

Tim Hobson, Lydia France, Sam Greenbury, Luke Hare, Pamela Wochner

专题命中 安全评测 :trustworthy(title)

Comments 10 pages, 4 figures, presented at the International Conference on AI and the Digital Economy (CADE 2023), Venice, Italy. Replaces the preprint version, with minor changes & additions based on reviewers' comments

Journal ref International Conference on AI and the Digital Economy (CADE 2023), 2023, pp. 31-40

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.14880 2023-09-28 cs.CV 71%

SGAligner : 3D Scene Alignment with Scene Graphs

Sayan Deb Sarkar, Ondrej Miksik, Marc Pollefeys, Daniel Barath, Iro Armeni

专题命中 安全评测 :alignment(title)

Comments Accepted at ICCV 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.12132 2023-06-22 cs.SE 71%

ChatGPT as a tool for User Story Quality Evaluation: Trustworthy Out of the Box?

Krishna Ronanki, Beatriz Cabrero-Daniel, Christian Berger

专题命中 安全评测 :trustworthy(title)

Comments 9 Pages, 2 Tables, 1 Figure. Accepted at AI-Assisted Agile Software Development Workshop (Co-located with XP 2023)

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.03842 2023-05-09 cs.DB 71%

Data Station: Delegated, Trustworthy, and Auditable Computation to Enable Data-Sharing Consortia with a Data Escrow

Siyuan Xia, Zhiru Zhu, Chris Zhu, Jinjin Zhao, Kyle Chard, Aaron J. Elmore, Ian Foster, Michael Franklin, Sanjay Krishnan, Raul Castro Fernandez

专题命中 安全评测 :trustworthy(title)

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.07862 2023-04-07 cs.HC 71%

What Do Children and Parents Want and Perceive in Conversational Agents? Towards Transparent, Trustworthy, Democratized Agents

Jessica Van Brummelen, Maura Kelleher, Mingyan Claire Tian, Nghi Hoang Nguyen

专题命中 安全评测 :trustworthy(title)

Comments 18 pages, 9 figures, submitted to IDC 2023, for associated appendix: https://gist.github.com/jessvb/fa1d4c75910106d730d194ffd4d725d3

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.15000 2023-03-28 cs.NI 71%

Explanation-Guided Deep Reinforcement Learning for Trustworthy 6G RAN Slicing

Farhad Rezazadeh, Hatim Chergui, Josep Mangues-Bafalluy

专题命中 安全评测 :trustworthy(title)

Comments 6 Pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.05424 2022-08-30 cs.HC 71%

"If it didn't happen, why would I change my decision?": How Judges Respond to Counterfactual Explanations for the Public Safety Assessment

Yaniv Yacoby, Ben Green, Christopher L. Griffin, Finale Doshi Velez

专题命中 安全评测 :safety(title)

Comments Accepted at HCOMP '22 and at CHI'22 Workshop on Human-Centered Perspectives in Explainable AI (HCXAI)

详情

展开后加载摘要…

URL PDF HTML 收藏
2111.14566 2021-11-30 astro-ph.IM 71%

Building Trustworthy Machine Learning Models for Astronomy

Michelle Ntampaka, Matthew Ho, Brian Nord

专题命中 安全评测 :trustworthy(title)

Comments Prepared for the Astronomical Data Analysis Software and Systems (ADASS) XXXI Proceedings

详情

展开后加载摘要…

URL PDF HTML 收藏
2109.04165 2021-09-10 cs.HC 71%

Modelling GDPR-Compliant Explanations for Trustworthy AI

Francesco Sovrano, Fabio Vitali, Monica Palmirani

专题命中 安全评测 :trustworthy(title)

详情

展开后加载摘要…

URL PDF HTML 收藏
2008.01980 2020-08-06 cs.HC cs.CV 71%

More Than Accuracy: Towards Trustworthy Machine Learning Interfaces for Object Recognition

Hendrik Heuer, Andreas Breiter

专题命中 安全评测 :trustworthy(title)

Comments UMAP '20: Proceedings of the 28th ACM Conference on User Modeling, Adaptation and Personalization

Journal ref UMAP 2020: Proceedings of the 28th ACM Conference on User Modeling, Adaptation and Personalization

详情

展开后加载摘要…

URL PDF HTML 收藏
1810.11556 2020-02-25 cs.RO 71%

Efficient and Trustworthy Social Navigation Via Explicit and Implicit Robot-Human Communication

Yuhang Che, Allison M. Okamura, Dorsa Sadigh

专题命中 安全评测 :trustworthy(title)

Journal ref IEEETransactionsonRobotics,pp(99):1-16,2020

详情

展开后加载摘要…

URL PDF HTML 收藏
1812.05444 2018-12-14 cs.CR cs.LO 71%

Pluralize: a Trustworthy Framework for High-Level Smart Contract-Draft

Zaynah Dargaye, Antonella Pozzo, Sara Tucci-Piergiovanni

专题命中 安全评测 :trustworthy(title)

详情

展开后加载摘要…

URL PDF HTML 收藏
1805.10401 2018-05-29 cs.CR 71%

Unsupervised Learning for Trustworthy IoT

Nikhil Banerjee, Thanassis Giannetsos, Emmanouil Panaousis, Clive Cheong Took

专题命中 安全评测 :trustworthy(title)

Comments 9 pages, 9 figures, 2018 IEEE International Conference on Fuzzy Systems

详情

展开后加载摘要…

URL PDF HTML 收藏
1705.08180 2017-05-24 cs.CV 71%

Correlation Alignment by Riemannian Metric for Domain Adaptation

Pietro Morerio, Vittorio Murino

专题命中 安全评测 :alignment(title)

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.22510 2026-08-25 cs.AI 新提交 70%

ClawProBench: Trace-Aware Evaluation of AI Agents with Runtime Coverage and Frozen Workplace-Style Holdouts

ClawProBench:基于运行时覆盖与冻结工作空间保留集的轨迹感知AI智能体评估

YuanHang Xiao

机构 * The Chinese University of Hong Kong(香港中文大学)

专题命中 安全评测 :alignment(abstract);safety(abstract);分类 cs.AI

AI总结 该研究提出ClawProBench基准,基于OpenClaw运行时构建,通过含102场景的全配置集与68场景的冻结保留集评估AI智能体,采用安全门控公式评分,发现最终答案排行榜存在缺陷,不同评估视角的智能体排名差异显著。

Comments 29 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.21736 2026-08-25 cs.LG 新提交 70%

Adaptive Multilevel Twisted Sequential Monte Carlo for Rare Events Estimation in Language Models

用于语言模型稀有事件估计的自适应多级扭曲序贯蒙特卡洛方法

Zixuan Liu, Fangzheng Wu, Brian Summa, Zizhan Zheng

专题命中 安全评测 :alignment(abstract);safety(abstract);分类 cs.LG

AI总结 针对现有扭曲序贯蒙特卡洛稀有事件估计依赖稀有正样本导致不可靠的问题,提出自适应多级扭曲SMC,通过逐步稀有中间事件学习扭曲,提升语言模型稀有不安全行为概率估计准确性,助力模型安全评估与对齐。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.12792 2026-08-25 cs.CR cs.AI 版本更新 70%

Silent Alarm: A J-Space Protocol for Comparing Danger Recognition Across Models and Quantization Levels

无声警报:一种用于跨模型和量化级别比较危险识别的J空间协议

Roman Prosvirnin, Victor Minchenkov, Alexey Soldatov, Vladimir Bashun, Anton Sergeev

机构 * HSE University(俄罗斯高等经济大学)

专题命中 安全评测 :safety(abstract);jailbreak(abstract);分类 cs.AI

AI总结 研究提出JADR协议,通过J空间测量模型内部表示,不依赖外部评判模型,计算本地运行。应用于六个模型跨越三种权重表示形式,以SafetyAUC指标比较,能显著区分不同安全机制模型,捕捉量化差异。

Comments 17 pages, 12 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.00869 2026-08-25 cs.LG 版本更新 70%

Enhancing LLM Metacognition via Cognitive Pairwise Training

通过认知成对训练增强LLM元认知

Weitao Li, Hao Zhou, Xuanyu Lei, Fandong Meng, Yuanhang Liu, Jingyi Ren, Ante Wang, Xiaolong Wang, Yuanchi Zhang, Fuwen Luo, Guangwen Yang, Lin Gan, Weizhi Ma, Yang Liu

机构 * National Engineering Laboratory for Intelligent Information Processing, Academy of Mathematics and Physics, Chinese Academy of Sciences(智能信息处理国家工程实验室,中国科学院数学物理研究所) University of Science and Technology of China(中国科学技术大学)

专题命中 安全评测 :alignment(abstract);trustworthy(abstract);分类 cs.LG

AI总结 提出认知成对训练(CPT),通过成对比较推理轨迹来学习区分可靠与不可靠推理,从而提升LLM的推理与元认知权衡。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.20797 2026-08-24 cs.AI 新提交 70%

Automated Trajectory Evaluation for Mobile Agents via Step-Level Consequence Reasoning and Aggregation

基于步骤级结果推理与聚合的移动智能体自动轨迹评估

Pengshuai Yang, Zijing Gao, Xue Yu, Benhui Zhuang, Bo Yuan, Junlan Feng

机构 * Jiutian Research(九天研究院) China Mobile(中国移动)

专题命中 安全评测 :alignment(abstract);safety(abstract);分类 cs.AI

AI总结 本研究提出CRATE(含安全评估扩展版CRATE-S)这一VLM-as-judge框架,通过步骤级推理解决移动智能体轨迹评估的上下文过载与安全缺失问题,在AndroidWorld、MobileRisk数据集上取得优于现有方法的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏