OpenFOAMGPT 2.0: end-to-end, trustworthy automation for computational fluid dynamics
专题命中 安全评测 :trustworthy(title)
AI 大模型
大模型对齐、安全、越狱、红队、提示注入和可信评测。
专题命中 安全评测 :trustworthy(title)
专题命中 安全评测 :trustworthy(title)
Comments 16 pages, 3 figures
专题命中 安全评测 :alignment(title)
Comments Accepted to CVPR 2025
专题命中 安全评测 :alignment(title)
Comments Accepted by ICLR 2025
专题命中 安全评测 :alignment(title)
Comments IEEE Transactions on Pattern Analysis and Machine Intelligence, Manuscript Info: 17 Pages, 13 Figures, and 6 Tables
专题命中 安全评测 :trustworthy(title)
Comments Short Vision Paper, Accepted by ICSE'25-NIER
专题命中 安全评测 :alignment(title)
专题命中 安全评测 :alignment(title)
Comments 7 pages, 3 figures
专题命中 安全评测 :trustworthy(title)
专题命中 安全评测 :trustworthy(title)
专题命中 安全评测 :alignment(title)
Comments Accepted to NeurIPS 2023
专题命中 安全评测 :trustworthy(title)
Comments 10 pages, 4 figures, presented at the International Conference on AI and the Digital Economy (CADE 2023), Venice, Italy. Replaces the preprint version, with minor changes & additions based on reviewers' comments
Journal ref International Conference on AI and the Digital Economy (CADE 2023), 2023, pp. 31-40
专题命中 安全评测 :alignment(title)
Comments Accepted at ICCV 2023
专题命中 安全评测 :trustworthy(title)
Comments 9 Pages, 2 Tables, 1 Figure. Accepted at AI-Assisted Agile Software Development Workshop (Co-located with XP 2023)
专题命中 安全评测 :trustworthy(title)
专题命中 安全评测 :trustworthy(title)
Comments 18 pages, 9 figures, submitted to IDC 2023, for associated appendix: https://gist.github.com/jessvb/fa1d4c75910106d730d194ffd4d725d3
专题命中 安全评测 :trustworthy(title)
Comments 6 Pages, 6 figures
专题命中 安全评测 :safety(title)
Comments Accepted at HCOMP '22 and at CHI'22 Workshop on Human-Centered Perspectives in Explainable AI (HCXAI)
专题命中 安全评测 :trustworthy(title)
Comments Prepared for the Astronomical Data Analysis Software and Systems (ADASS) XXXI Proceedings
专题命中 安全评测 :trustworthy(title)
专题命中 安全评测 :trustworthy(title)
Comments UMAP '20: Proceedings of the 28th ACM Conference on User Modeling, Adaptation and Personalization
Journal ref UMAP 2020: Proceedings of the 28th ACM Conference on User Modeling, Adaptation and Personalization
专题命中 安全评测 :trustworthy(title)
Journal ref IEEETransactionsonRobotics,pp(99):1-16,2020
专题命中 安全评测 :trustworthy(title)
专题命中 安全评测 :trustworthy(title)
Comments 9 pages, 9 figures, 2018 IEEE International Conference on Fuzzy Systems
专题命中 安全评测 :alignment(title)
ClawProBench:基于运行时覆盖与冻结工作空间保留集的轨迹感知AI智能体评估
机构 * The Chinese University of Hong Kong(香港中文大学)
专题命中 安全评测 :alignment(abstract);safety(abstract);分类 cs.AI
AI总结 该研究提出ClawProBench基准,基于OpenClaw运行时构建,通过含102场景的全配置集与68场景的冻结保留集评估AI智能体,采用安全门控公式评分,发现最终答案排行榜存在缺陷,不同评估视角的智能体排名差异显著。
Comments 29 pages, 4 figures
用于语言模型稀有事件估计的自适应多级扭曲序贯蒙特卡洛方法
专题命中 安全评测 :alignment(abstract);safety(abstract);分类 cs.LG
AI总结 针对现有扭曲序贯蒙特卡洛稀有事件估计依赖稀有正样本导致不可靠的问题,提出自适应多级扭曲SMC,通过逐步稀有中间事件学习扭曲,提升语言模型稀有不安全行为概率估计准确性,助力模型安全评估与对齐。
无声警报:一种用于跨模型和量化级别比较危险识别的J空间协议
机构 * HSE University(俄罗斯高等经济大学)
专题命中 安全评测 :safety(abstract);jailbreak(abstract);分类 cs.AI
AI总结 研究提出JADR协议,通过J空间测量模型内部表示,不依赖外部评判模型,计算本地运行。应用于六个模型跨越三种权重表示形式,以SafetyAUC指标比较,能显著区分不同安全机制模型,捕捉量化差异。
Comments 17 pages, 12 figures
通过认知成对训练增强LLM元认知
机构 * National Engineering Laboratory for Intelligent Information Processing, Academy of Mathematics and Physics, Chinese Academy of Sciences(智能信息处理国家工程实验室,中国科学院数学物理研究所) ; University of Science and Technology of China(中国科学技术大学)
专题命中 安全评测 :alignment(abstract);trustworthy(abstract);分类 cs.LG
AI总结 提出认知成对训练(CPT),通过成对比较推理轨迹来学习区分可靠与不可靠推理,从而提升LLM的推理与元认知权衡。
基于步骤级结果推理与聚合的移动智能体自动轨迹评估
机构 * Jiutian Research(九天研究院) ; China Mobile(中国移动)
专题命中 安全评测 :alignment(abstract);safety(abstract);分类 cs.AI
AI总结 本研究提出CRATE(含安全评估扩展版CRATE-S)这一VLM-as-judge框架,通过步骤级推理解决移动智能体轨迹评估的上下文过载与安全缺失问题,在AndroidWorld、MobileRisk数据集上取得优于现有方法的性能。