Multi-word Units in Machine Translation: why the Tip of the Iceberg Remains Problematic – and a Tentative Corpus-driven Solution

(2019) MUMTT2019, The 4th Workshop on Multi-word Units in Machine Translation and Translation Technology — Location: University of Málaga (27.September.2019)

Files

MUMTT19JPColson.pdf
  • Restricted Access
  • Adobe PDF
  • 557.78 KB

Details

Authors
Abstract
Neural machine translation (NMT) has recently made considerable progress in the improvement of the quality of the produced texts. Among the new features of NMT are the fluidity of the translations and their successful treatment of multi-word units. In this paper we first report the results of a global and automated evaluation of the percentage of phraseology in the translations produced by Google Translate and DeepL. A corpus-based approach makes it possible to estimate that both NMT systems succeed in producing an average percentage of phraseology that is quite natural and may sometimes even be higher than in natural language production by native speakers. Taking a closer look at some problematic cases, however, reveals that the phraseological value of NMT systems may be deceptive, as they are often inable to cope with contextual complexity and low-frequency idioms.
Affiliations

Citations

Colson, J.-P. (2019). Multi-word Units in Machine Translation: why the Tip of the Iceberg Remains Problematic – and a Tentative Corpus-driven Solution. MUMTT2019, The 4th Workshop on Multi-word Units in Machine Translation and Translation Technology, University of Málaga. https://hdl.handle.net/2078.5/123403