arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.31632cs.LG

EEGAgentBench:在短时与长时程脑电图分析上对LLM智能体的基准测试

EEGAgentBench: Benchmarking LLM Agents on Short- and Long-Horizon EEG Analysis

Huyu Wu, Weining Weng, Yuchen Liu, Yiqiang Chen, Yang Gu

AI总结:

提出EEGAgentBench统一基准,覆盖短时与长时程EEG分析,含6种应用和10个工具,评估29个LLM智能体,揭示其在长时程任务中证据积累与多步推理的不足。

AI中文摘要:

脑电图(EEG)分析正从短片段分类向长时程解释演变,后者要求迭代证据积累、多步推理以及专用信号处理工具的协调使用。尽管大型语言模型(LLM)近期在作为自主智能体进行EEG分析方面展现出潜力,但现有的EEG智能体评估仍然零散,覆盖的任务有限,时间跨度狭窄,协议不一致,且未对智能体的推理、工具使用和工作流构建能力提供全面评估。为弥补这一空白,我们提出EEGAgentBench,一个用于系统评估LLM智能体在短时与长时程EEG分析上的统一基准。EEGAgentBench涵盖从知识问答到睡眠分期的六种代表性EEG应用,信号时长从2秒到近23小时不等,预测目标从类别标签到事件区间和epoch级序列。该设计支持在知识推理、短时程解释、长时程事件检测和序列理解上进行统一评估。基准进一步提供10个确定性EEG分析工具,仅暴露与任务相关的信号测量值。因此,智能体必须自主选择工具、迭代积累证据并构建多步工作流。在评估中,我们对来自15个模型家族的29个前沿LLM进行了基准测试。结果表明,EEGAgentBench能有效区分超越模型规模和推理成本的智能体能力,同时揭示了当前LLM智能体在长时程EEG分析中的显著局限,尤其是在持续证据积累和多步推理方面。

英文摘要:

Electroencephalography (EEG) analysis is evolving from short-segment classification toward long-horizon interpretation that demands iterative evidence accumulation, multi-step reasoning, and coordinated use of specialized signal-processing tools. Although large language models (LLMs) have recently shown promise as autonomous agents for EEG analysis, existing EEG agentic evaluations remain fragmented, covering limited tasks over narrow temporal horizons with inconsistent protocols, and providing no comprehensive assessment of agents' reasoning, tool-use, and workflow construction capabilities. To address this gap, we propose \textbf{EEGAgentBench}, a unified benchmark for systematically evaluating LLM agents on short- and long-horizon EEG analysis. EEGAgentBench spans six representative EEG applications ranging from knowledge question answering to sleep staging. It encompasses signal durations from 2 seconds to nearly 23 hours, with prediction targets ranging from class labels to event intervals and epoch-level sequences. This design supports unified evaluation across knowledge reasoning, short-horizon interpretation, long-horizon event detection, and sequential understanding. The benchmark further provides 10 deterministic EEG analysis tools that expose only task-relevant signal measurements. Agents must therefore select tools autonomously, accumulate evidence iteratively, and construct multi-step workflows. For evaluation, we benchmark 29 frontier LLMs from 15 model families. Results demonstrate that EEGAgentBench effectively distinguishes agent capabilities beyond model scale and inference cost, while revealing substantial limitations of current LLM agents in long-horizon EEG analysis, particularly in sustained evidence accumulation and multi-step reasoning.

↑