arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2504.03541cs.CL

分解程序与对齐话语后的多样化上下文示例选择改进了语义解析

Diverse In-Context Example Selection After Decomposing Programs and Aligned Utterances Improves Semantic Parsing

  • Indian Institute of Technology Bombay(印度理工学院孟买分校)
  • Amazon(亚马逊公司)

机构由 AI 辅助整理,请以论文原文为准。

Mayank Kothyari, Sunita Sarawagi, Soumen Chakrabarti, Gaurav Arora, Srujana Merugu

更新

AI总结:

提出SCUD4ICL系统,通过分解AST树片段并映射对齐话语,结合多样化上下文示例选择,显著提升LLMs在语义解析任务中的准确率。

AI中文摘要:

LLMs越来越多地被用作从自然语言话语到结构化程序的seq2seq翻译器,这一过程被称为语义解释。与原子标签或token序列不同,程序自然地表示为抽象语法树(ASTs)。这种结构化表示引发了与呈现给LLM的上下文示例的设计和选择相关的新问题。我们专注于将可用的ICE树池分解为片段,其中一些片段可能更适合于解决测试实例。接下来,我们提出如何使用LLM(额外的调用)结合提示的语法约束来自动地将这些片段映射到对应的话语。最后,我们适应并扩展了一种最近的方法,用于多样化的ICE选择,以处理完整和片段化的ICE实例。我们在流行的多样化语义解析基准上评估我们的系统SCUD4ICL,展示了我们提出的分解多样化演示方法带来的明显准确率提升。对于较小的LLMs、具有更大标注树的ICE池以及较低资源语言的程序,益处尤为显著。

英文摘要:

LLMs are increasingly used as seq2seq translators from natural language utterances to structured programs, a process called semantic interpretation. Unlike atomic labels or token sequences, programs are naturally represented as abstract syntax trees (ASTs). Such structured representation raises novel issues related to the design and selection of in-context examples (ICEs) presented to the LLM. We focus on decomposing the pool of available ICE trees into fragments, some of which may be better suited to solving the test instance. Next, we propose how to use (additional invocations of) an LLM with prompted syntax constraints to automatically map the fragments to corresponding utterances. Finally, we adapt and extend a recent method for diverse ICE selection to work with whole and fragmented ICE instances. We evaluate our system, SCUD4ICL, on popular diverse semantic parsing benchmarks, showing visible accuracy gains from our proposed decomposed diverse demonstration method. Benefits are particularly notable for smaller LLMs, ICE pools having larger labeled trees, and programs in lower resource languages.

补充信息

↑