arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

LLM智能体作为计算类型学家

LLM Agents as Computational Typologists

Changbing Yang, Christopher Hammerly, Freda Shi, Jian Zhu

arXiv 2609.07791首次发表:更新:

发表机构

University of British Columbia; University of Waterloo; Vector Institute(不列颠哥伦比亚大学; 滑铁卢大学; 矢量研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出AUTOTYPOLOGIST,一种基于LLM的智能体,通过检索语法章节、分析IGT并采用ReAct工作流,实现基于证据的类型学分析,在特征编码和假设检验任务中展现潜力,但仍需专家验证。

AI 中文摘要

语言类型学依赖于专家对跨语言参考语法的分析,这使得大规模跨语言比较既费力又难以扩展。我们提出了AUTOTYPOLOGIST,一个用于基于证据的参考语法类型学分析的LLM智能体。该智能体能够检索相关的语法章节,分析行间注音文本(IGT),并使用ReAct风格的工作流对类型学假设进行迭代推理。我们在类型学特征编码任务上,针对专家标注评估了该系统,并在类型学假设检验任务中,使用25个开源参考语法对类型学共性进行了评估。在类型学特征编码中,在不同的信息约束下,该智能体能够综合参考语法散文中的信息,但在仅有目标语言IGT的情况下仍面临挑战。在类型学假设检验中,该智能体能够综合跨语言证据,并识别支持案例和反例。这些发现表明,LLM智能体可以支持可扩展且可检查的类型学分析,但仍需专家验证。

英文摘要

Linguistic typology relies on expert analysis of reference grammars across languages, making large-scale crosslinguistic comparison labor-intensive and unscalable. We introduce AUTOTYPOLOGIST, an LLM agent for evidence-grounded typological analysis over reference grammars. The agent is capable of retrieving relevant grammar sections, analyzing interlinear glossed text (IGT), and iteratively reasoning over typological hypotheses using a ReAct-style workflow. We evaluate the system on TYPOLOGICAL FEATURE CODING against expert annotations and TYPOLOGICAL HYPOTHESIS TESTING with typological universals using 25 open-source reference grammars. Operating under different information constraints in TYPOLOGICAL FEATURE CODING, the agent can synthesize information from reference grammar prose but still faces challenges with only IGTs in the target language. In TYPOLOGICAL HYPOTHESIS TESTING, the agent can synthesize crosslinguistic evidence and identify both supporting cases and counterexamples. These findings suggest that LLM agents can support scalable and inspectable typological analysis, while still requiring expert validation.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑