AI 中文总结
研究针对大语言模型实际行为难测、现有评估协议有局限的问题,提出StabilityBench基准测试工具,通过注入现实模拟增强现有基准,应用于四个基准测试评估九个模型,发现性能不稳定,还提出StabilityBench-Mini以实现更现实评估。
AI 中文摘要
人工智能助手越来越多地应用于医疗或政府服务等高风险场景,但由于强烈的上下文依赖性,其实际行为仍知之甚少。当前评估协议遵循深度防御范式,从传统基准测试到实时或对抗性测试。此类基准测试大多是静态单轮的,限制了捕捉对话场景中现实世界变异性的能力。我们提出了StabilityBench,它将单轮基准查询转换为多轮交互历史。通过注入现实用户模拟增强现有基准测试,同时保留原始任务意图。我们将其应用于四个基准测试,评估九个大语言模型,结果显示模型性能不稳定,凸显了静态评估的局限性。为此,我们提出了StabilityBench-Mini,在不增加成本的情况下实现更现实的评估。
英文摘要
AI Assistants are increasingly deployed in high-stakes settings, such as healthcare or government services. Yet their real-world behavior remains poorly understood due to strong context dependence. Current evaluation protocols follow a defense-in-depth paradigm with compounding layers of safeguards, ranging from traditional benchmarks to live or adversarial testing. Such benchmarks remain largely static and single-turn, limiting their ability to capture real-world variability in conversational settings. We propose StabilityBench, a principled, general and model-agnostic benchmark operator that transforms single-turn benchmark queries into multi-turn interaction histories. StabilityBench augments existing benchmarks by injecting realistic user simulations, through demographic proxies or sycophantic baits, while preserving original task intent. We apply StabilityBench to four benchmarks spanning mathematical reasoning, health question-answering and safety, and evaluate nine large language models under these conditions. Our results show that model performance is consistently unstable under these injections, with considerable performance degradations on three out of four benchmarks studied. These highlight important limitations of static evaluations and motivate more realistic evaluation settings. To this end, we propose StabilityBench-Mini: a size-preserving variant of StabilityBench that samples across diversification axes, enabling more realistic evaluation without increasing costs.