可信大型语言模型、智能体AI与多模态系统的统一评估框架
A Unified Evaluation Framework for Trustworthy Large Language Models, Agentic AI, and Multimodal Systems
浏览论文内容
中文总结 AI 辅助
针对现有AI可信评估碎片化问题,提出统一框架,通过八个维度整合LLM、智能体及多模态系统评估,并引入元评估与安全覆盖机制,确保评估结果可解释且可靠。
中文摘要 AI 辅助
仅凭基准测试分数无法全面评估现代人工智能系统的可信度。大型语言模型(LLM)、智能体系统和多模态模型(MLLM)需要不同形式的评估,但其评估证据必须保持可解释性,以支持开发和监督。我们提出了一个统一框架,通过八个可信度维度(能力、鲁棒性、安全性、公平性、透明度、治理、监督和效率)连接输出级、轨迹级和跨模态评估。该框架保留系统特定指标,同时将原生测量映射到通用性能区间,并附带不确定性估计和可追溯证据。元评估层检验评估本身的有效性、可靠性和可复现性。多维画像揭示系统优缺点,而安全关键覆盖机制防止聚合分数掩盖关键失败。与治理框架、国际标准和欧盟监管要求的映射将技术评估与监督需求联系起来。该框架为评估系统性能及其支撑证据的可信度提供了结构化基础,跨部署场景的实证验证仍是下一步的关键工作。
英文摘要
Benchmark scores alone provide an incomplete basis for assessing the trustworthiness of modern artificial intelligence systems. Large language models (LLMs), agentic systems, and multimodal models (MLLMs) require different forms of assessment, yet their evaluation evidence must remain interpretable for development and oversight. We propose a unified framework that connects output-level, trajectory-level, and cross-modal assessment through eight trustworthiness dimensions: capability, robustness, safety, fairness, transparency, governance, oversight, and efficiency. The framework preserves system-specific metrics while mapping native measurements to common performance bands, accompanied by uncertainty estimates and traceable evidence. A meta-evaluation layer examines the validity, reliability, and reproducibility of the evaluation itself. Multidimensional profiles expose strengths and weaknesses, while safety-critical overrides prevent aggregate scores from masking critical failures. Mappings to governance frameworks, international standards, and European Union regulatory requirements connect technical assessment with oversight needs. The framework provides a structured basis for assessing both system performance and the credibility of the evidence supporting it, with empirical validation across deployment contexts remaining an essential next step.
发表机构
- Vector Institute for Artificial Intelligence(向量人工智能研究所)
机构由 AI 辅助整理,请以论文原文为准。