arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.04156cs.AIcs.LG

BrainBench:面向全面脑电理解的大语言模型基准测试

BrainBench: Benchmarking Large Language Models for Comprehensive EEG Understanding

Yangxuan Zhou, Yuning Chen, Chen Wu, Jiquan Wang, Shijian Li, Gang Pan, Sha Zhao

首次发表
浏览论文内容

中文总结 AI 辅助

该研究推出BrainBench基准测试,涵盖4个子集共17个数据集等,评估大语言模型在两种范式下的脑电理解能力,为相关研究提供可复现测试平台。

中文摘要 AI 辅助

脑电图(EEG)分析不仅限于为记录分配预定义标签,还需要连接自然语言指令、信号处理、定量证据和科学解释的工作流程,我们将这种能力称为“全面脑电理解”。然而,现有评估主要针对孤立的解码任务或特定系统演示,未能充分量化大语言模型(LLM)的能力。我们推出 BrainBench,这是一个用于全面、指令条件下脑电理解的统一基准。它包含四个子集——基础分析、睡眠评估、神经认知评估和生理整合——涵盖17个数据集、numcases个任务和超过numinstances个真实数据实例。给定一条指令和带有可选生理信号的脑电记录,系统必须执行分析并生成一份科学依据充分的报告,必要时生成人工制品。输出通过数值、分类、集合、序列、语义和人工制品验证进行评估。我们在两种范式下对nummodels个代表性LLM进行了超过10万次执行的评估:通过CodeAct的自主代码执行和通过BrainAgent的结构化智能体分析。结果在模型、子集、难度级别和执行范式之间存在显著差异,表明脑电理解能力取决于模型及其操作方式。BrainBench为推进基于LLM的脑电理解提供了可复现的测试平台。代码和基准将很快发布,评估结果将持续更新。

英文摘要

Electroencephalography (EEG) analysis extends beyond assigning predefined labels to recordings; it requires workflows connecting natural-language instructions, signal processing, quantitative evidence, and scientific interpretation. We term this capability \emph{comprehensive EEG understanding}. Existing evaluations, however, primarily target isolated decoding tasks or system-specific demonstrations, leaving the competence of large language models (LLMs) insufficiently quantified. We introduce \benchmarkname{}, a unified benchmark for comprehensive, instruction-conditioned EEG understanding. It comprises four subsets---Foundational Analysis, Sleep Assessment, Neurocognitive Assessment, and Physiological Integration---covering 17 datasets, \numcases{} tasks, and over \numinstances{} real-data instances. Given an instruction and EEG recordings with optional physiological signals, a system must perform the analysis and produce a scientifically grounded report and, when required, artifacts. Outputs are assessed through numerical, categorical, set, sequence, semantic, and artifact validation. We evaluate \nummodels{} representative LLMs across more than 100K executions under two paradigms: autonomous code execution with CodeAct and structured agentic analysis with BrainAgent. Results vary substantially across models, subsets, difficulty levels, and execution paradigms, showing that EEG competence depends on the model and its operationalization. \benchmarkname{} provides a reproducible testbed for advancing LLM-based EEG understanding. The code and benchmark will be released soon, with evaluation results continuously updated.

发表机构

  • Zhejiang University(浙江大学)
  • College of Computer Science and Technology, Zhejiang University(浙江大学计算机科学与技术学院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑