Crosslingual Embeddings are Essential in UNMT for Distant Languages: An English to IndoAryan Case Study

Banerjee, Tamali; Murthy V, Rudra; Bhattacharyya, Pushpak

Computer Science > Computation and Language

arXiv:2106.04995 (cs)

[Submitted on 9 Jun 2021]

Title:Crosslingual Embeddings are Essential in UNMT for Distant Languages: An English to IndoAryan Case Study

Authors:Tamali Banerjee, Rudra Murthy V, Pushpak Bhattacharyya

View PDF

Abstract:Recent advances in Unsupervised Neural Machine Translation (UNMT) have minimized the gap between supervised and unsupervised machine translation performance for closely related language pairs. However, the situation is very different for distant language pairs. Lack of lexical overlap and low syntactic similarities such as between English and Indo-Aryan languages leads to poor translation quality in existing UNMT systems. In this paper, we show that initializing the embedding layer of UNMT models with cross-lingual embeddings shows significant improvements in BLEU score over existing approaches with embeddings randomly initialized. Further, static embeddings (freezing the embedding layer weights) lead to better gains compared to updating the embedding layer weights during training (non-static). We experimented using Masked Sequence to Sequence (MASS) and Denoising Autoencoder (DAE) UNMT approaches for three distant language pairs. The proposed cross-lingual embedding initialization yields BLEU score improvement of as much as ten times over the baseline for English-Hindi, English-Bengali, and English-Gujarati. Our analysis shows the importance of cross-lingual embedding, comparisons between approaches, and the scope of improvements in these systems.

Subjects:	Computation and Language (cs.CL); Machine Learning (cs.LG)
Cite as:	arXiv:2106.04995 [cs.CL]
	(or arXiv:2106.04995v1 [cs.CL] for this version)
	https://doi.org/10.48550/arXiv.2106.04995

Submission history

From: Tamali Banerjee [view email]
[v1] Wed, 9 Jun 2021 11:31:27 UTC (2,112 KB)

Computer Science > Computation and Language

Title:Crosslingual Embeddings are Essential in UNMT for Distant Languages: An English to IndoAryan Case Study

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computation and Language

Title:Crosslingual Embeddings are Essential in UNMT for Distant Languages: An English to IndoAryan Case Study

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators