arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2026-01-13 至 2026-01-13 共收录 32 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 32 篇

2601.06757 2026-01-13 cs.CL cs.AI 82%

MTMCS-Bench: Evaluating Contextual Safety of Multimodal Large Language Models in Multi-Turn Dialogues

MTMCS-Bench: 多轮对话中多模态大语言模型上下文安全性的评估

Zheyuan Liu, Dongwhi Kim, Yixin Wan, Xiangchi Yuan, Zhaoxuan Tan, Fengran Mo, Meng Jiang

机构 * University of Notre Dame(诺丁汉大学) University of California, Los Angeles(加州大学洛杉矶分校) Georgia Institute of Technology(佐治亚理工学院) University of Montreal(蒙特利尔大学)

专题命中 安全评测 :safety(title,abstract);分类 cs.CL、cs.AI

AI总结 MTMCS-Bench评估多模态大语言模型在多轮对话中的上下文安全性,揭示了安全与效用之间的权衡及现有防护措施的不足。

Comments A benchmark of realistic images and multi-turn conversations that evaluates contextual safety in MLLMs under two complementary settings

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11218 2026-01-13 cs.CL cs.AI 81%

The Curious Case of Factual (Mis)Alignment between LLMs' Short- and Long-Form Answers

事实性(误解)一致性在LLM短形式和长形式回答中的奇特案例

Saad Obaid ul Islam, Anne Lauscher, Goran Glavaš

机构 * WüNLP, CAIDAS, University of Würzburg(乌尔姆大学) Data Science Group, University of Hamburg(汉堡大学)

专题命中 安全评测 :alignment(title,abstract);分类 cs.CL、cs.AI

AI总结 本文提出SLAQ框架,揭示LLM在短长形式回答中存在事实一致性问题,通过机制分析发现重叠内部结构与78%预测准确率,挑战现有评估实践。

Comments Code: https://github.com/WorldHellow/SLAQ/tree/main

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.06704 2026-01-13 cs.LG cs.AI cs.CV 81%

Beyond Perfect Scores: Proof-by-Contradiction for Trustworthy Machine Learning

超越完美分数:基于矛盾证明的可信机器学习证明

Dushan N. Wadduwage, Dineth Jayakody, Leonidas Zimianitis

机构 * Old Dominion University(旧 Dominion 大学)

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI、cs.LG

AI总结 本文提出基于矛盾证明的可信度测试方法,通过随机排列标签来评估模型是否依赖真实临床线索,从而提高机器学习在生物医学中的可信度和临床应用潜力。

Comments 13 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.06167 2026-01-13 cs.LG cs.AI 81%

Parent-Guided Adaptive Reliability (PGAR): A Behavioural Meta-Learning Framework for Stable and Trustworthy AI

指导式自适应可靠性(PGAR):一种用于稳定和可信赖AI的行为元学习框架

Anshum Rankawat

机构 * Anshum Rankawat(独立研究者)

专题命中 安全评测 :trustworthy(title);safety(abstract);分类 cs.AI、cs.LG

AI总结 PGAR通过引入监督父层提升AI的稳定性与可靠性,通过有界可靠性指数调节学习率,实现更稳健的优化与学习过程。

Comments 9 pages, 8 figures, 2 tables. Submitted to IEEE Transactions on Artificial Intelligence

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09195 2026-01-13 cs.CV 78%

Towards Trustworthy Dermatology MLLMs: A Benchmark and Multimodal Evaluator for Diagnostic Narratives

迈向可信的皮肤科多模态大语言模型:一种基准和多模态评估器用于诊断叙述

Yuhao Shen, Jiahe Qian, Shuping Zhang, Zhangtianyi Chen, Tao Lu, Juexiao Zhou

机构 * School of Data Science, The Chinese University of Hong Kong, Shenzhen(数据科学学院,香港中文大学(深圳)) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) Department of Dermatology, The First Affiliated Hospital, Shantou University Medical College(皮肤科,汕头大学医学院第一附属医院)

专题命中 安全评测 :trustworthy(title);alignment(abstract)

AI总结 本文提出DermBench和DermEval用于评估皮肤科多模态大语言模型的诊断叙述,通过结合基准和自动评估器,实现临床意义的可重复评估,验证模型性能与专家评分的一致性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.07239 2026-01-13 cs.AI 70%

Stochastic CHAOS: Why Deterministic Inference Kills, and Distributional Variability Is the Heartbeat of Artifical Cognition

随机CHAOS:为何确定性推断会杀死,而分布性变化是人工认知的心跳

Tanmay Joshi, Shourya Aggarwal, Anusa Saha, Aadi Pandey, Shreyash Dhoot, Vighnesh Rai, Raxit Goswami, Aman Chadha, Vinija Jain, Amitava Das

机构 * Raapid Lab, USA(美国Raapid实验室) Apple, USA(美国苹果公司) Google, USA(美国谷歌公司) Pragya Lab, BITS Pilani Goa, India(印度BITS Pilani Goa大学Pragya实验室)

专题命中 安全评测 :alignment(abstract);safety(abstract);分类 cs.AI

AI总结 本文主张随机CHAOS,认为确定性推断会压制LLM的不确定性建模和安全对齐,强调分布性变化对人工认知的重要性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.02109 2026-01-13 cs.AI cs.CL cs.CY 67%

Deep Value Benchmark: Measuring Whether Models Generalize Deep Values or Shallow Preferences

深度价值基准:衡量模型是泛化深度价值还是浅层偏好

Joshua Ashkinaze, Hua Shen, Saipranav Avula, Eric Gilbert, Ceren Budak

机构 * University of Michigan Ann Arbor(密歇根大学安娜堡分校) New York University Shanghai(纽约大学上海)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI、cs.CY

AI总结 深度价值基准通过测试模型泛化深度价值还是浅层偏好,评估AI对齐的核心能力。

Comments NeurIPS 2025 (Spotlight)

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.06464 2026-01-13 cs.CV 67%

On the Adversarial Robustness of 3D Large Vision-Language Models

关于3D大视觉-语言模型的对抗鲁棒性

Chao Liu, Ngai-Man Cheung

机构 * Singapore University of Technology and Design (SUTD)(新加坡科技设计大学)

专题命中 安全评测 :alignment(abstract);safety(abstract)

AI总结 本文研究了3D大视觉-语言模型的对抗鲁棒性,提出两种攻击策略评估其鲁棒性,发现其在无目标攻击下较脆弱,但在有目标攻击中更稳健。

Comments Under Review

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.06193 2026-01-13 cs.LG cs.AI cs.CL 67%

MLB: A Scenario-Driven Benchmark for Evaluating Large Language Models in Clinical Applications

MLB:面向临床应用的大型语言模型评估场景驱动基准

Qing He, Dongsheng Bi, Jianrong Lu, Minghui Yang, Zixiao Chen, Jiacheng Lu, Jing Chen, Nannan Du, Xiao Cu, Sijing Wu, Peng Xiang, Yinyin Hu, Yi Guo, Chunpu Li, Shaoyang Li, Zhuo Dong, Ming Jiang, Shuai Guo, Liyun Feng, Jin Peng, Jian Wang, Jinjie Gu, Junwei Liu

机构 * Ant Group(蚂蚁集团) Zhejiang University(浙江大学) Health Information Center of Zhejiang Province(浙江省健康信息中心) Peking University(北京大学)

专题命中 安全评测 :safety(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 MLB基准通过场景驱动的评估方法,评估大型语言模型在临床应用中的表现,揭示模型在结构化任务与患者交互场景间的性能差异。

Comments 11 pages, 4 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.06106 2026-01-13 cs.LG cs.AI cs.CL cs.CV cs.MA 67%

Judge Model for Large-scale Multimodality Benchmarks

大规模多模态基准的判断模型

Min-Han Shih, Yu-Hsin Wu, Yu-Wei Chen

机构 * Department of Electrical and Computer Engineering, Viterbi School of Engineering, University of Southern California(电气与计算机工程系,维特比工程学院,南加州大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本文提出了一种多模态判断模型,用于评估多种任务,通过聚合多模态判断并生成诊断反馈,展示了其在多模态AI研究中的应用潜力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.03113 2026-01-13 cs.CE 67%

A Probabilistic Digital Twin of UK En Route Airspace for Training and Evaluating AI Agents for Air Traffic Control

基于英国航路空域的概率数字孪生:用于训练和评估空管AI代理

Nick Pepper, Adam Keane, Amy Hodgkin, Dewi Gould, Edward Henderson, Lynge Lauritsen, Christos Vlahos, George De Ath, Richard Everson, Richard Cannon, Alvaro Sierra Castro, John Korna, Ben Carvell, Marc Thomas

专题命中 安全评测 :safety(abstract);trustworthy(abstract)

AI总结 本文提出基于英国航路空域的概率数字孪生,用于训练和评估空管AI代理,通过结合历史与实时数据及概率模型,提供高保真虚拟环境以支持AI代理的开发与评估。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.07767 2026-01-13 cs.LG cs.CL 62%

Are LLM Decisions Faithful to Verbal Confidence?

大型语言模型的决策是否忠实于其语言自信度?

Jiawei Wang, Yanfei Zhou, Siddartha Devic, Deqing Fu

机构 * University of Southern California(南加州大学)

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL、cs.LG

AI总结 研究发现大型语言模型在高惩罚条件下缺乏战略决策能力,其语言自信度与实际决策不一致,影响AI系统的可信度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.07313 2026-01-13 cs.LG cs.AI 62%

Explaining Machine Learning Predictive Models through Conditional Expectation Methods

通过条件期望方法解释机器学习预测模型

Silvia Ruiz-España, Laura Arnal, François Signol, Juan-Carlos Perez-Cortes, Joaquim Arlandis

机构 * ITI, Universitat Politècnica de València(ITI,瓦伦西亚理工大学)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG

AI总结 本文提出MUCE方法,通过多变量条件期望捕捉预测变化,结合稳定性与不确定性指标,提升模型局部可解释性与可信度。

Comments 24 pages, 15 figures. Silvia Ruiz-España and Laura Arnal contributed equally to this work

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.07226 2026-01-13 cs.AI cs.CL 62%

Lost in the Noise: How Reasoning Models Fail with Contextual Distractors

陷入噪声中:推理模型在上下文干扰中的失败

Seongyun Lee, Yongrae Jo, Minju Seo, Moontae Lee, Minjoon Seo

机构 * University of Illinois Chicago(伊利诺伊大学芝加哥分校)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

AI总结 本文提出RARE方法,通过激励识别噪声中的有用信息,提升模型在上下文干扰下的鲁棒性,揭示增加计算量反而导致噪声环境下性能下降的反比例趋势。

Comments Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.07651 2026-01-13 cs.LG cs.CY 62%

Enhancing Binary Encoded Crime Linkage Analysis Using Siamese Network

利用孪生网络增强二进制编码的犯罪关联分析

Yicheng Zhan, Fahim Ahmed, Amy Burrell, Matthew J. Tonkin, Sarah Galambos, Jessica Woodhams, Dalal Alrajeh

专题命中 安全评测 :safety(abstract);分类 cs.CY、cs.LG

AI总结 本文提出利用孪生网络增强二进制编码的犯罪关联分析,通过学习潜在表示和整合地理时间特征,提升关联准确性并提供可解释的见解。

Comments AAAI 2026, 7 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.05623 2026-01-13 cs.SE cs.AI cs.CL 62%

Deployability-Centric Infrastructure-as-Code Generation: Fail, Learn, Refine, and Succeed through LLM-Empowered DevOps Simulation

以部署为中心的基础设施即代码生成:通过LLM赋能的DevOps模拟实现失败、学习、细化和成功

Tianyi Zhang, Shidong Pan, Zejun Zhang, Zhenchang Xing, Xiaoyu Sun

机构 * Australian National University Canberra Australia New York University \& Columbia University USA Nanyang Technological University Singapore Australian National University Australia Australian National University New York University \& Columbia University Nanyang Technological University

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

AI总结 本文提出IaCGen框架,通过迭代反馈机制提升IaC模板的部署性,实验表明其在10次迭代内可使模板部署成功率提升至91.6%,并进一步通过人工反馈将性能提升至90%以上。

Comments Accepted by FSE 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.06750 2026-01-13 cs.CV cs.AI cs.CL 62%

Benchmarking Egocentric Clinical Intent Understanding Capability for Medical Multimodal Large Language Models

医疗多模态大语言模型的视点临床意图理解能力基准测试

Shaonan Liu, Guo Yu, Xiaoling Luo, Shiyi Zheng, Wenting Chen, Jie Liu, Linlin Shen

机构 * Shenzhen University(深圳大学) Stanford University(斯坦福大学) City University of Hong Kong(香港城市大学)

专题命中 安全评测 :safety(abstract);分类 cs.CL、cs.AI

AI总结 本文提出MedGaze-Bench,首个评估医疗多模态大语言模型视点临床意图理解能力的基准测试,通过三维意图框架和陷阱QA机制,揭示现有模型在手术、急救和诊断任务中对意图理解的不足。

Comments 16 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.07329 2026-01-13 cs.CL 57%

BayesRAG: Probabilistic Mutual Evidence Corroboration for Multimodal Retrieval-Augmented Generation

BayesRAG: 基于概率互证的多模态检索增强生成

Xuan Li, Yining Wang, Haocai Luo, Shengping Liu, Jerry Liang, Ying Fu, Weihuang, Jun Yu, Junnan Zhu

机构 * University of Science and Technology of China(中国科学技术大学) Unisound AI Technology Co.Ltd(Unisound AI技术有限公司) MAIS, Institute of Automation, Chinese Academy of Sciences(自动化研究所,中国科学院)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

AI总结 BayesRAG通过概率互证机制提升多模态检索增强生成的性能,有效解决异构模态的隔离问题。

Comments 17 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.07274 2026-01-13 cs.CL 57%

Towards Comprehensive Semantic Speech Embeddings for Chinese Dialects

面向中文方言的全面语义语音嵌入

Kalvin Chang, Yiwen Shao, Jiahong Li, Dong Yu

机构 * Tencent AI Labs(腾讯AI实验室) Tencent(腾讯)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

AI总结 本文提出通过仅使用ASR数据训练语音编码器,实现中文方言与普通话的跨方言语义对齐,并在新基准测试中展示了语音到语音检索和ASR性能的突破。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.07004 2026-01-13 cs.CR cs.AI 57%

MemTrust: A Zero-Trust Architecture for Unified AI Memory System

MemTrust:面向统一AI内存系统的零信任架构

Xing Zhou, Dmitrii Ustiugov, Haoxin Shang, Kisson Lin

机构 * Independent Researcher(独立研究者) Supermem AI Inc.(Supermem AI公司)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

AI总结 MemTrust提出了一种五层零信任架构,旨在实现AI内存系统的本地等效安全性和高效协作能力,通过TEE保护和跨应用共享协议提升系统可信度。

Comments 18 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.06944 2026-01-13 cs.CV cs.AI 57%

SketchJudge: A Diagnostic Benchmark for Grading Hand-drawn Diagrams with Multimodal Large Language Models

SketchJudge: 一种用于用多模态大语言模型评分手绘图的诊断基准

Yuhang Su, Mei Wang, Yaoyao Zhong, Guozhang Li, Shixing Li, Yihan Feng, Hua Huang

机构 * School of Artificial Intelligence, Beijing Normal University(人工智能学院,北京师范大学)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

AI总结 SketchJudge是一个用于评估多模态大语言模型在评分手绘图任务中诊断能力的基准,通过1015份手绘学生回答验证了当前视觉-语言对齐在符号和嘈杂环境中的脆弱性。

Comments 8 pages for the main text (excluding references and the limitations section); 37 pages in total including appendices

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14317 2026-01-13 cs.LG 57%

Intervention Efficiency and Perturbation Validation Framework: Capacity-Aware and Robust Clinical Model Selection under the Rashomon Effect

干预效率与扰动验证框架:在拉索莫恩效应下考虑容量的临床模型选择

Yuwen Zhang, Viet Tran, Paul Weng

专题命中 安全评测 :trustworthy(abstract);分类 cs.LG

AI总结 本文提出干预效率与扰动验证框架,用于在资源约束下稳健选择临床模型,解决拉索莫恩效应下的模型选择问题。

Comments Accepted to the Workshop on Navigating Model Uncertainty and the Rashomon Effect: From Theory and Tools to Applications and Impact (AAAI 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.08728 2026-01-13 cs.CR cs.AI cs.IR 57%

Securing RAG: A Risk Assessment and Mitigation Framework

保障RAG:风险评估与缓解框架

Lukas Ammann, Sara Ott, Christoph R. Landolt, Marco P. Lehmann

机构 * Eastern Switzerland University of Applied Sciences (OST)(东瑞士应用科学大学(OST)) Cyber-Defence Campus, armasuisse Science and Technology(网络安全校园,armasuisse科技)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

AI总结 本文提出了一种结合RAG特定安全考虑与通用安全指南的框架,以保障RAG系统的安全性和可信度。

Comments 8 pages, 3 figures, Sara Ott and Lukas Ammann contributed equally. This work has been submitted to the IEEE for possible publication

Journal ref 2025 IEEE Swiss Conference on Data Science (SDS)

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.06845 2026-01-13 cs.AI 57%

Code Evolution for Control: Synthesizing Policies via LLM-Driven Evolutionary Search

代码进化用于控制:通过LLM驱动的进化搜索合成策略

Ping Guo, Chao Li, Yinglan Feng, Chaoning Zhang

机构 * Shenzhen Yuyi Tech. Co. Ltd.(深圳优衣科技有限公司) City University of Hong Kong(香港城市大学) Shenzhen University(深圳大学) University of Electronic Science and Technology of China(电子科技大学)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

AI总结 本文提出利用LLM驱动的进化搜索生成可解释的控制策略,通过代码进化方法合成可执行代码,提升自主系统控制策略的可解释性和可靠性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.06516 2026-01-13 cs.HC cs.LG eess.SP 57%

Pareto-Optimal Model Selection for Low-Cost, Single-Lead EMG Control in Embedded Systems

低成本单通道EMG控制的帕累托最优模型选择

Carl Vincent Ladres Kho

机构 * Minerva University(明维大学)

专题命中 安全评测 :safety(abstract);分类 cs.LG

AI总结 本文提出了一种低成本单通道EMG控制的帕累托最优模型选择方法,通过评估多种模型架构,确定随机森林在嵌入式控制中的最优性能。

Comments 15 pages main text, 51 pages total including appendices. 18 figures. Code and dataset available at: https://github.com/CarlKho-Minerva/v2-emg-muscle

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.06141 2026-01-13 cs.CY 57%

An LLM -Powered Assessment Retrieval-Augmented Generation (RAG) For Higher Education

基于大语言模型的评估检索增强生成(RAG)系统用于高等教育

Reza Vatankhah Barenji, Nazila Salimi, Sina Khoshgoftar

专题命中 安全评测 :alignment(abstract);分类 cs.CY

AI总结 本研究提出基于RAG架构的大语言模型驱动评估系统,通过整合评分标准和范文生成高质量反馈,实现大规模、一致的评估反馈,提升学生自主学习能力。

Comments 19 Pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.05774 2026-01-13 cs.AI 57%

ConstraintLLM: A Neuro-Symbolic Framework for Industrial-Level Constraint Programming

ConstraintLLM: 一种用于工业级约束编程的神经符号框架

Weichun Shi, Minghao Liu, Wanting Zhang, Langchen Shi, Fuqi Jia, Feifei Ma, Jian Zhang

机构 * Hangzhou Institute for Advanced Study, UCAS, Hangzhou, China(杭州高等研究院,UCAS,杭州,中国) University of Oxford, Oxford, UK(牛津大学,牛津,英国) University of Science and Technology Beijing, Beijing, China(北京科技大学,北京,中国) SKLCS and Key Laboratory of System Software, ISCAS, Beijing, China(SKLCS和系统软件重点实验室,ISCAS,北京,中国) Laboratory of Parallel Software and Computational Science, ISCAS, Beijing, China(并行软件与计算科学实验室,ISCAS,北京,中国) University of Chinese Academy of Sciences, Beijing, China(中国科学院大学,北京,中国)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

AI总结 ConstraintLLM是一种专为约束编程设计的神经符号框架,通过引入Constraint-Aware Retrieval Module和Tree-of-Thoughts框架,实现了在工业级约束编程基准上的高性能求解。

Comments Accepted to the 2025 Conference on Empirical Methods in Natural Language Processing (EMNLP 2025), Main Conference

Journal ref Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, pages 15999-16019

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11890 2026-01-13 cs.RO cs.AI 57%

Integrating Symbolic RL Planning into a BDI-based Autonomous UAV Framework: System Integration and SIL Validation

将符号RL规划整合到基于BDI的自主无人机框架中:系统集成与SIL验证

Sangwoo Jeon, Juchul Shin, YeonJe Cho, Gyeong-Tae Kim, Seongwoo Kim

专题命中 安全评测 :safety(abstract);分类 cs.AI

AI总结 本研究提出AMAD-SRL框架,整合符号RL与BDI方法,通过SIL验证提升无人机任务效率75%。

Comments This submission has been withdrawn by the authors due to institutional and contractual requirements related to security and export-control review

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.07186 2026-01-13 cs.RO 50%

PROTEA: Securing Robot Task Planning and Execution

PROTEA:保障机器人任务规划与执行的安全性

Zainab Altaweel, Mohaiminul Al Nahian, Jake Juettner, Adnan Siraj Rakin, Shiqi Zhang

专题命中 安全评测 :safety(abstract)

AI总结 PROTEA通过LLM-as-a-Judge机制评估机器人任务计划的安全性,解决维度和历史挑战,提升任务规划系统的鲁棒性和安全性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.06779 2026-01-13 cs.CR 50%

CyberLLM-FINDS 2025: Instruction-Tuned Fine-tuning of Domain-Specific LLMs with Retrieval-Augmented Generation and Graph Integration for MITRE Evaluation

CyberLLM-FINDS 2025:基于检索增强生成和图集成的领域特定LLM指令微调方法用于MITRE评估

Vasanth Iyer, Leonardo Bobadilla, S. S. Iyengar

专题命中 安全评测 :alignment(abstract)

AI总结 本文提出了一种基于检索增强生成和图集成的领域特定LLM微调方法,通过STIX威胁情报实现与MITRE ATT&CK技术的对齐,提升网络安全威胁情报分析的准确性。

Comments 12 pages

详情

展开后加载摘要…

URL PDF HTML 收藏