AI 中文总结
该研究推出PredAct-Bench基准,以教育为测试平台评估带噪声工具的对话智能体,扩展信任校准指标并评估13种LLMs,发现当前模型无法在工具噪声下为教师提供足够可见性,助力构建更优AI决策支持系统。
AI 中文摘要
大型语言模型(LLMs)正越来越多地被部署到任务导向型对话系统中,这类系统支持在教育、医疗、金融等高风险领域开展多步骤决策。然而,现有基准通常假设工具输出完全准确,却忽略了实际部署的系统必须在工具存在噪声、人类决策者对智能体的信任本身不确定的情况下运行,这类情况在实践中十分常见,例如临床医生使用诊断预测工具,或顾问依赖基于历史记录预测学生成绩的模型。我们推出PredAct-Bench,一个用于评估搭配统计上不完善工具的对话智能体的基准,以教育作为可衡量的测试平台,该平台具备真实结果和明确的干预决策。首先,我们构建了一个用于AI辅助人类决策的基准,其中AI使用带噪声的预测器来辅助指导用户。其次,我们引入了回合级别的相对AI依赖(RAIR)和相对自我依赖(RSR)指标,将现有的信任校准框架扩展至多轮对话。第三,我们在两个教育数据集上评估了13种最先进的闭源和开源LLMs,这两个数据集分别是OULAD(英国开放大学的真实评估轨迹)和PredAct-CS(包含60门课程的真实期末成绩结果与合成生成的每周成绩轨迹),同时还开展了一项针对教师和助教的人类研究。我们发现,当工具存在噪声时,最先进的模型应向教师提供可见性,以防止他们过度依赖错误建议或幻觉,但当前模型未能做到这一点。我们提供PredAct-Bench,以帮助构建更好的LLMs作为AI决策支持系统来辅助教师。
英文摘要
Large Language Models (LLMs) are increasingly deployed in task-oriented dialogue systems that support multi-step decision-making in high-stakes domains such as education, healthcare, and finance. However, existing benchmarks typically assume perfectly accurate tool outputs, overlooking the reality that deployed systems must operate with noisy tools and human decision-makers whose trust in the agent is itself uncertain. Such conditions are common in practice, for example, a clinician using a diagnostic prediction tool or an advisor relying on a model that forecasts student outcomes from historical records. We introduce PREDACTBENCH, a benchmark for evaluating dialogue agents paired with statistically imperfect tools, using education as a measurable testbed where ground truth outcomes and clear intervention decisions are available. First, we build a benchmark for AI-assisted human decision-making, where the AI uses noisy predictors to help guide a user. Second, we introduce episode-level Relative AI-Reliance (RAIR) and Relative self-reliance (RSR) metrics, extending prior trust calibration framework to multi-turn dialogue. Third, we evaluate 13 state-of-the-art closed and open source LLMs on two educational datasets, OULAD (real assessment trajectories from the UK Open University) and PREDACT-CS (60 courses with real final grade outcomes and synthetically generated weekly score trajectories), alongside a human study with instructors and teaching assistants. We find that when tools are noisy, SOTA models are supposed to provide visibility to teachers so that they do not over-rely on wrong suggestions or hallucinations, but current models fail to do that. We offer PREDACTBENCH to help build better LLMs as AI decision support systems to help teachers.