arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.13721cs.CLquant-ph

通过自动预群超标注扩展印地语量子自然语言处理

Scaling Hindi Quantum Natural Language Processing through Automatic Pregroup Supertagging

Gautami Sanjay Naik, Krishna Bhatia, Mithun Paul Saint-Germain, H Aswath Babu

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出将印地语预群超标注作为词元级分类任务,利用380句人工标注语料评估多种方法,发现上下文回退达64.56%最佳准确率,证明自动分配可行且能减少人工标注依赖。

中文摘要 AI 辅助

量子自然语言处理(QNLP)利用预群语法将语法结构转化为图表示和量子电路。近期印地语QNLP研究表明,印地语特有的预群语法可以支持语法敏感的组成模型,但语法类型分配仍主要依赖人工,限制了可扩展性。本文将自动印地语预群超标注形式化为一个词元级分类任务。使用一个包含380个印地语句子的人工标注语料库,我们评估了词汇、上下文、基于提示、词汇修复以及后缀/形态感知方法。结果表明,在这种低资源设置下,简单的词汇和上下文模型表现强劲:上下文回退实现了64.56%的最佳完成准确率,而原始Qwen2.5提示仅达到11.65%。词汇修复将LLM辅助预测提升至64.08%,展示了用符号语法知识约束生成输出的价值。诊断分析进一步表明,已见和无歧义词元比未见词元容易得多,后缀/形态特征提高了karaka词元准确率但未提升整体性能。这些结果表明,自动印地语预群分配是可行的,并可在未来多语言QNLP流程中减少对人工标注的依赖。

英文摘要

Quantum Natural Language Processing (QNLP) uses pregroup grammars to translate grammatical structure into diagrammatic representations and quantum circuits. Recent Hindi QNLP work has shown that Hindi-specific pregroup grammars can support grammar-sensitive compositional models, but grammatical type assignment is still largely manual, limiting scalability. This paper formulates automatic Hindi pregroup supertagging as a token-level classification task. Using a manually annotated corpus of 380 Hindi sentences, we evaluate lexical, contextual, prompting-based, lexical-repair, and suffix/morphology-aware methods. Results show that simple lexical and contextual models are strong in this low-resource setting: contextual backoff achieves the best completed accuracy of 64.56\%, while raw Qwen2.5 prompting reaches only 11.65\%. Lexical repair raises LLM-assisted prediction to 64.08\%, demonstrating the value of constraining generative outputs with symbolic grammar knowledge. Diagnostic analysis further shows that seen and unambiguous tokens are much easier than unseen tokens, and suffix/morphology features improve karaka-token accuracy but not overall performance. These results show that automatic Hindi pregroup assignment is feasible and can reduce reliance on manual annotation in future multilingual QNLP pipelines.

发表机构

  • Indian Institute of Information Technology Dharwad(印度达尔瓦德信息技术学院)
  • Fractal Analytics(Fractal Analytics公司)
  • Arizona State University(亚利桑那州立大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑