用于可复现全模态基础模型评估的可组合评估系统
A Composable Evaluation System for Reproducible Omni-Modal Foundation Model Evaluation
浏览论文内容
中文总结 AI 辅助
针对全模态基础模型评估工具互不兼容的问题,提出OmniEvaluator系统,通过单一接口整合现有资源,支持复现与跨模型比较,成本低于商业LLM评判器。
中文摘要 AI 辅助
构建全模态基础模型需要对其在文本、图像、视频和音频领域进行评估。各模态均有优秀的评估工具包,但它们的推理引擎、提示约定和指标实现互不兼容,导致从业者需为每个工具链维护独立环境,且仍难以跨工具链比较结果。OmniEvaluator 正是源于我们自身模型开发的这一需求:它不重新实现基准,而是在更高层面连接现有推理引擎和精选评估库,通过单一接口提供4种推理后端、4种评估框架及超过1000个基准。每次运行都记录为捕获完整配置的制品,以实现精确复现,结果会汇入共享仪表板用于跨模型比较。联邦模式可在并发评估间共享GPU推理服务器,内置的验证器体积小到可在CPU上运行,能在规则评分因配置不匹配随引擎和提示波动时保持分数稳定,其效果与高成本商业LLM评判器相当,却无重复API成本。该系统、演示视频和仪表板均已公开(此https链接)。
英文摘要
Building an omni-modal foundation model means evaluating it across text, image, video, and audio. Excellent evaluation toolkits exist for each modality, but their inference engines, prompt conventions, and metric implementations are mutually incompatible, so practitioners end up maintaining separate environments for every toolchain and still struggle to compare results across them. OmniEvaluator grew out of this need in our own model development: rather than reimplementing benchmarks, it connects existing inference engines and curated evaluation libraries at a higher level, exposing four inference backends, four evaluation frameworks, and over a thousand benchmarks through a single interface. Every run is recorded as an artifact capturing the full configuration for exact reproduction, and results flow into a shared dashboard for cross-model comparison. A federated mode shares GPU inference servers across concurrent evaluations, and a built-in verifier, small enough to run on CPU, keeps its score stable across engines and prompts where rule-based scoring fluctuates under configuration mismatch, matching cost-efficient commercial LLM judges without their recurring API cost. The system, demo video, and dashboard are publicly available. (https://github.com/naver-ai/omni-evaluator)
发表机构
- NAVER Cloud AI(NAVER 云AI)
- Korea University(高丽大学)
- KAIST AI(韩国科学技术院AI)
- Seoul National University(首尔大学)
机构由 AI 辅助整理,请以论文原文为准。