AI 中文总结
RADAR是一个可审计运行时,通过保留证据的探索、类型化操作符记录和运行时验证,使智能体分析选择可检查并支持修订,实现大声失败,在KramaBench等基准上显著提升性能。
AI 中文摘要
大型语言模型(LLMs)使得数据科学智能体能够对异构文件执行多步骤分析。然而,关于数据源、范围或统计定义的不正确选择常常导致静默错误:计算成功执行,但产生看似合理实则错误的输出,未能回答预期问题。为缓解这一问题,我们提出了RADAR,一个可审计的运行时,使智能体的分析选择可检查,并通过执行反馈支持其修订。RADAR通过三个核心机制运作。首先,一个保留证据的探索模块检索与任务相关的内容,同时保留源位置和观察覆盖率。接下来,运行时使用类型化操作符记录智能体声明的输入、操作参数和结果观察。最后,运行时验证检查提议的操作与这些观察。当检测到冲突时,运行时拒绝操作或提供诊断反馈,允许智能体在错误传播之前修订其选择。这种设计使智能体能够大声失败,同时将语义解释留给LLM。在KramaBench上,RADAR在完整源检索下达到0.723的总体分数,在提供金源下达到0.747,相对于最强基线分别有35.9%和28.8%的相对增益。在KramaBench之外,RADAR在DA-Code上实现了14.0%的相对性能增益,在DABStep上实现了59.3%的相对增益,展示了其在多样化智能体数据分析工作流中的适用性。
英文摘要
Large language models (LLMs) have enabled data-science agents to automate multi-step analyses over heterogeneous files. However, incorrect choices regarding data sources, scope, or statistical definitions often lead to silent errors: computations execute successfully but produce plausible yet incorrect outputs that fail to answer the intended question. To mitigate this, we present RADAR, an auditable runtime that makes an agent's analytical choices inspectable and supports their revision through execution feedback. RADAR operates through three core mechanisms. First, an evidence-preserving exploration module retrieves task-relevant content while retaining source locations and observation coverage. Next, the runtime uses typed operators to record the agent's declared inputs, operation arguments, and resulting observations. Finally, runtime validation checks proposed operations against these observations. When a conflict is detected, the runtime rejects the operation or provides diagnostic feedback, allowing the agent to revise its choices before errors propagate. This design enables agents to fail loudly while leaving semantic interpretation to the LLM. On KramaBench, RADAR achieves overall scores of 0.723 with full source retrieval and 0.747 with gold sources supplied, corresponding to relative gains of 35.9% and 28.8% over the strongest baselines. Beyond KramaBench, RADAR achieves relative performance gains of 14.0% on DA-Code and 59.3% on DABStep, demonstrating its applicability across diverse agentic data-analysis workflows.
Comments12 pages, 5 figures