AI 中文总结
本研究提出一种结合层次化认知过程建模与过程监督的框架,利用LLM、LoRA和MoE提升场景安全理解的可解释性和性能,并构建了带过程标签的新数据集。
AI 中文摘要
场景安全理解在多个关键领域的情境感知中扮演着生死攸关的角色。传统方法依赖于学习场景与安全等级之间的直接映射,往往缺乏可解释性,限制了其在关键应用中的可靠性。克服这一挑战的有效途径在于解读人类认知过程,并为机器模型赋予类似的认知能力。本研究探索了一种将场景安全认知过程建模与过程监督相结合的有效方法。具体而言,我们首先构建了一个层次化的认知安全结构,这促使我们基于多步推理和过程标签开发了一个新颖、高质量的场景安全理解数据集。该数据集既作为基准,也作为提升大型语言模型(LLMs)安全推理能力的资源,同时通过信息流和基于显著性的技术,能够对中间推理步骤进行细粒度分析。在此基础上,我们引入了一个模块化且灵活的过程监督框架,该框架反映了人类认知的层次性。该框架以LLMs为核心架构,并结合了低秩自适应(LoRA)和混合专家(MoE)策略,以实现专家模块的专业化和协作,每个模块负责整体推理链中的特定子过程。系统的实验评估和分析证实,与传统方法相比,我们的框架展现出优越的可解释性和性能特征。
英文摘要
Scene safety understanding plays a life-or-death role in situational awareness in various critical domains. Traditional methods that rely on learning direct mappings between scenes and safety levels often lack interpretability, limiting their reliability in critical applications. An effective approach to overcoming this challenge lies in interpreting human cognitive processes and equipping machine models with analogous cognitive capabilities. This work explores an effective way of integrating scene safety cognitive process modeling and process supervision. Specifically, we first construct a hierarchical cognitive safety structure, which motivates the development of a novel, high-quality scene safety understanding dataset based on multi-step reasoning with process labels. This dataset serves both as a benchmark and a resource to improve the safety reasoning capabilities of Large Language Models (LLMs), while also enabling a granular analysis of intermediate reasoning steps through information flow and saliency-based techniques. Building upon this foundation, we introduce a modular and flexible process supervision framework that reflects the hierarchical nature of human cognition. This framework leverages LLMs as the core architecture and incorporates Low-Rank Adaptation(LoRA) and Mixture-of-Experts (MoE) strategies to enable specialization and collaboration among expert modules, each tasked with specific sub-processes of the overall reasoning chain. Systematic experimental evaluations and analyses confirm that our framework exhibits superior interpretability and performance characteristics compared to traditional approaches.