立场:停止每次都被动地修补模型,开始主动的测试驱动人工智能开发
Position: Stop Reactively Patching Your Model Every Time and Start Proactive Test-Driven AI Development
浏览论文内容
中文总结 AI 辅助
针对现代AI系统多样用例下的维护问题,提出用主动测试驱动飞轮解决反应式飞轮的局限性,通过创建“测试空间”映射反馈数据到任务目标,经数学证明主动飞轮能以更少迭代实现更好长期扩展性。
中文摘要 AI 辅助
许多现代人工智能系统旨在在多样、开放的用例下运行。为帮助泛化已部署系统,许多维护管道使用反应式人工智能飞轮,根据用户行为反馈修补模型。但作为主要维护机制时,它常忽略问题 broader context,无法预防未来边缘情况,导致更多不必要迭代。而且由于开放世界用例的长尾特性,收集剩余错误在统计上越来越难。本文主张需主动测试驱动飞轮解决其局限性,通过创建“测试空间”将反馈数据映射到任务目标,使飞轮从被动变为主动。并通过数学证明主动飞轮比被动飞轮能以更少迭代实现更好的长期扩展性。
英文摘要
Many modern AI systems are designed to operate under diverse, open-ended, use-cases. To help generalize deployed systems, many deployed-system maintenance pipelines use a reactive AI flywheel that observes emerging feedback from user behavior (errors) and patches the model accordingly. However, when used as the primary maintenance mechanism, these flywheels often ignore the broader context of these errors within the system's objectives, failing to preempt potential future edge cases, which leads to more unnecessary flywheel iterations. Also, it is statistically increasingly difficult to collect remaining errors due to the long-tail nature of open-world use-cases. This position paper argues that a proactive test-driven flywheel is required to address reactive flywheel's limitations and to approach a generalizable system. We advocate for creating a "test space" to technically map feedback data to task objectives, evolving the flywheel from reactive to proactive. We augment our position by mathematically proving a proactive one achieves better long-term scaling with fewer iterations than the reactive flywheel.