发表机构
The George Washington University; California Institute of Technology; University of California, Berkeley(乔治华盛顿大学; 加州理工学院; 加州大学伯克利分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对ATC仍以人工为主的问题,设计5种提示结构并在9种LLMs上实验,发现简洁提示表现最佳,注入正确历史可修复脚本严格提示的错误,明确了LLM辅助ATC的路径与局限。
AI 中文摘要
空中交通管制(ATC)通信是一项安全关键型对话,尽管空中交通管理的其他部分已实现半自动化,但该领域仍主要由人工驱动。本文通过实验评估大语言模型(LLMs)是否能生成符合操作实际的ATC传输内容。一条飞越旧金山“海湾游览”航线的通用航空飞行被手动转录并用作基准事实(P0)。通过飞行员在环流程,我们设计了5种约束程度递增的提示结构(P1-P5),并将其嵌入一个有状态多轮管道中,其中模型扮演ATC角色,与固定的飞行员转录文本对话,同时基于不断积累的对话历史进行条件生成。我们在9种开源和闭源LLMs上,对提示、是否提供来自另一实验飞行的已完成转录文本作为上下文示例、模型是否基于自身先前回复或注入的基准事实历史进行条件生成这几个变量进行了控制。对话轮次通过词汇、结构和语义相似度指标,以及经人类专家注释验证的LLM作为评判者(GPT-5.5)进行评分。提供已完成示例可提升相似度,但收紧提示则无效:最简洁的提示表现最佳,而脚本最严格的提示会因自身错误在对话中累积而失效,注入正确历史可修复该问题。这些结果勾勒出LLM辅助ATC的具体路径及其当前局限。
英文摘要
Air traffic control (ATC) communication is a safety-critical dialogue that remains largely human-driven even as other parts of air traffic management have been semi-automated. In this article, we experimentally evaluate whether large language models (LLMs) can generate operationally realistic ATC transmissions. An experimental general-aviation flight flying over the San Francisco "Bay Tour" route is hand-transcribed and used as ground truth (P0). Through a pilot-in-the-loop process we design five prompt structures (P1-P5) of increasing constraint and embed them in a stateful multi-turn pipeline, where the model plays ATC to a fixed pilot transcript while conditioning on the accumulating dialogue history. Across nine open- and closed-source LLMs we vary the prompt, the presence of a worked transcript from a different experimental flight as an in-context example, and whether the model conditions on its own prior replies or on injected ground-truth history. Turns are scored with lexical, structural, and semantic similarity metrics and by an LLM-as-judge (GPT-5.5) validated against human expert annotation. Supplying a worked example improves similarity, but tightening the prompt does not: the lightest prompts perform best and the most heavily scripted one collapses as its own errors accumulate through the dialogue, which injecting correct history repairs. These results outline a concrete path and its current limits toward LLM-assisted ATC.
Comments39 pages, 12 figures, 7 tables