ARIA — 用于信息娱乐系统自主测试的智能体框架
ARIA - An Agentic Framework for Autonomous Testing of Infotainment Systems
- Critical Techworks, Portugal(Critical Techworks)
- Faculty of Engineering, University of Porto, Portugal(波尔图大学工程学院)
- LIACC, Faculty of Engineering, University of Porto, Portugal(波尔图大学工程学院LIACC)
- INESC TEC, Faculty of Engineering, University of Porto, Portugal(波尔图大学工程学院INESC TEC)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出多智能体LLM框架ARIA,针对安卓信息娱乐系统自主开展端到端视觉测试,经30个场景评估,其缺陷检测准确率高,证实了多智能体设计在工业级信息娱乐测试中的价值。
AI中文摘要:
汽车信息娱乐系统的验证仍依赖手动测试,该方法速度慢、成本高,且与敏捷发布和OTA更新不兼容。脚本自动化仅能提供部分帮助:它将测试逻辑与实现耦合,导致套件脆弱且维护成本高。现有的大语言模型(LLM)驱动框架大多针对网页或移动应用,采用单智能体或双智能体设置,将感知、规划、动作选择和验证等任务同时交由一两个模型承担,考虑到信息娱乐系统的复杂性,这类框架易出现幻觉和无效探索循环。我们提出ARIA(Autonomous Real-time Infotainment Assessment,自主实时信息娱乐评估),这是一种多智能体LLM框架,可通过视觉交互在安卓信息娱乐系统上自主运行端到端测试,采用每步包含四个专门智能体加一个报告阶段的闭环流水线。从单句场景(路径、动作、预期结果)出发,ARIA运行交互并生成报告、可复现脚本及每步的视觉证据。在某厂商的实体安卓信息娱乐系统上针对30个场景进行评估,ARIA完成了28个(93.3%)并给出判定结果(2个出错),其中20个(71.4%)与真实情况相符;它捕获了全部5个已知缺陷,无故障被判定为正常;其8个误报源于导航/图像限制和不支持的手势,表明多智能体LLM可在工业层面运行信息娱乐测试,同时揭示了低误报容忍度的代价。单智能体基线实验证实了多智能体设计的价值:在更强模型重访缩小差距前的首次运行中,单智能体的误报率远高于多智能体(72.0% vs. 52.6%),它将导航难度与系统故障混为一谈。我们报告了首次运行/重访后的结果、每个场景的 token/调用/成本,并通过重复运行表明稳定性与复杂性相关,故障检测完全一致,指向视觉测试的CI集成。
英文摘要:
Automotive infotainment validation still relies on manual testing, slow, costly, and incompatible with agile releases and OTA updates. Scripted automation only partly helps: it couples test logic to implementation, yielding brittle, high-maintenance suites. Existing LLM-driven frameworks mostly target web/mobile apps, using single- or dual-agent setups that overload one or two models with perception, planning, action selection, and validation at once, prone to hallucinations and unproductive exploration loops given infotainment complexity. We present ARIA (Autonomous Real-time Infotainment Assessment), a multi-agent LLM framework that autonomously runs end-to-end tests on Android infotainment systems via visual interaction, using a closed-loop pipeline of four specialized agents per step plus a report stage. From single-sentence scenarios (path, action, expected outcome), ARIA runs the interactions and produces reports, reproducible scripts, and visual evidence per step. Evaluated on a manufacturer's physical Android infotainment system across 30 scenarios, ARIA completed 28 (93.3%) with a verdict (2 errored), 20 of which (71.4%) matched ground truth. It caught all 5 known defects, no fault passed as working; its 8 false positives stem from navigation/image limits and unsupported gestures, showing multi-agent LLMs can run infotainment tests industrially while exposing the cost of a low false-positive tolerance. A single-agent baseline confirms the multi-agent design's value: on the first pass, before stronger-model revisitation narrows the gap, it shows a far higher false-positive rate (72.0% vs. 52.6%), conflating navigational difficulty with system failure. We report first-pass/post-revisitation results, token/call/cost per scenario, and show via repeated runs that stability tracks complexity, with fault detection perfectly consistent, pointing to CI integration of visual testing.