arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 1802.01433cs.CLcs.AIcs.LG

Interactive Grounded Language Acquisition and Generalization in a 2D World

发表机构百度研究院 · 深度学习技术及应用国家工程实验室
查看机构详情
  • Baidu Research(百度研究院)
  • National Engineering Laboratory for Deep Learning Technology and Applications(深度学习技术及应用国家工程实验室)

机构由 AI 辅助整理,请以论文原文为准。

Haonan Yu, Haichao Zhang, Wei Xu

首次发表 更新
浏览论文内容

英文摘要

We build a virtual agent for learning language in a 2D maze-like world. The agent sees images of the surrounding environment, listens to a virtual teacher, and takes actions to receive rewards. It interactively learns the teacher's language from scratch based on two language use cases: sentence-directed navigation and question answering. It learns simultaneously the visual representations of the world, the language, and the action control. By disentangling language grounding from other computational routines and sharing a concept detection function between language grounding and prediction, the agent reliably interpolates and extrapolates to interpret sentences that contain new word combinations or new words missing from training sentences. The new words are transferred from the answers of language prediction. Such a language ability is trained and evaluated on a population of over 1.6 million distinct sentences consisting of 119 object words, 8 color words, 9 spatial-relation words, and 50 grammatical words. The proposed model significantly outperforms five comparison methods for interpreting zero-shot sentences. In addition, we demonstrate human-interpretable intermediate outputs of the model in the appendix.

补充信息

↑