Multilingual Pixel Representations for Translation and Effective Cross-lingual Transfer

Pretrained multilingual translation models using either pixel or subword (bpe) representations trained on the many-to-one parallel TED-59 dataset, accompanying the EMNLP'23 paper "Multilingual Pixel Representations for Translation and Effective Cross-lingual Transfer." Models can be interacted with on the command line or through a script similarly to other fairseq models, but require our code extension for rendered text with pixel representations. Each model zip file contains: the fairseq model checkpoint, vocab files, language list file, and relevant sentencepiece model(s). We additionally package the TED-59 data here in raw extracted format for ease of comparison (original dataset release and paper by Qi et. al 2018). For more information, see our: Paper describing the method and training data [arXiv] Code repository with scripts [github]

Related Organizations

Microsoft (United States)
United States
Johns Hopkins University
United States

Impact byBIP!

	selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).	0
	popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.	Average
	influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).	Average
	impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.	Average

Found an issue? Give us feedback

0

Average