Skip to content

Objectives

aoliverg edited this page Sep 20, 2022 · 1 revision

Our project's main objective is to design, train and evaluate NMT systems between the Romance languages of the Iberian Peninsula. This objective will be achieved through the following specific objectives:

  • Compile parallel and monolingual corpora for the languages included in the proposal, paying special attention to the languages ​​with fewer resources: Asturian, Aragonese and Aranese.
  • Explore new techniques for training neural machine translation engines applied to the following Romance languages ​​of the Iberian Peninsula: Spanish, Portuguese, Catalan, Galician, Asturian, Aragonese and Aranese.
  • Train neural machine translation systems between Spanish and the rest of the languages of TAN-IBE, in both directions. This implies engines for the pairs: Spanish-Portuguese, Portuguese-Spanish, Spanish-Catalan, Catalan-Spanish, Spanish-Galician, Galician-Spanish, * Spanish-Asturian, Asturian-Spanish, Spanish-Aragonese, Aragonese-Spanish, Spanish-Aranese and Aranese-Spanish.
  • Train neural multilingual systems able to translate from and to all the languages ​​of the project.
  • Evaluate all the trained systems using automatic evaluation metrics and compare them, when possible, with existing machine translation systems.
  • Perform manual evaluations of at least the following resulting machine translation systems: Spanish-Asturian, Spanish-Aragonese and Spanish-Aranese. This manual evaluation will be performed measuring the post-edition effort.
  • Create guides and scripts that facilitate the training of neural machine translation engines in general, and more specifically for the language pairs of the project.
  • Publish the results of TAN-IBE with free licences. This includes the compiled corpora, the machine translation models and engines and the guides and scripts.

Clone this wiki locally