ALBERT — All Library Books, journals and Electronic Records Telegrafenberg

Treffer pro Seite

Treffer 1 - 1 | 1 Treffer

Alles auswählen Exportieren

Unbekannt

ParaMed: a parallel corpus for English–Chinese translation in the biomedical domain (2021)

Liu, Boxiang ; Huang, Liang

BioMed Central

In: BMC Medical Informatics and Decision Making. 2021; 21(1): 258. Published 2021 Sep 06. doi: 10.1186/s12911-021-01621-8.

zur Merkliste hinzufügen auf der Merkliste

Details

Publikationsdatum: 2021-09-06

Beschreibung: Background Biomedical language translation requires multi-lingual fluency as well as relevant domain knowledge. Such requirements make it challenging to train qualified translators and costly to generate high-quality translations. Machine translation represents an effective alternative, but accurate machine translation requires large amounts of in-domain data. While such datasets are abundant in general domains, they are less accessible in the biomedical domain. Chinese and English are two of the most widely spoken languages, yet to our knowledge, a parallel corpus does not exist for this language pair in the biomedical domain. Description We developed an effective pipeline to acquire and process an English-Chinese parallel corpus from the New England Journal of Medicine (NEJM). This corpus consists of about 100,000 sentence pairs and 3,000,000 tokens on each side. We showed that training on out-of-domain data and fine-tuning with as few as 4000 NEJM sentence pairs improve translation quality by 25.3 (13.4) BLEU for en$$ ightarrow$$ → zh (zh$$ ightarrow$$ → en) directions. Translation quality continues to improve at a slower pace on larger in-domain data subsets, with a total increase of 33.0 (24.3) BLEU for en$$ ightarrow$$ → zh (zh$$ ightarrow$$ → en) directions on the full dataset. Conclusions The code and data are available at https://github.com/boxiangliu/ParaMed.

Digitale ISSN: 1472-6947

Thema: Informatik , Medizin

Publiziert von BioMed Central

	Standort	Signatur	Erwartet	Verfügbarkeit

Andere fanden auch interessant ...

AKTUELLE ARTIKEL

S·F·X

Volltext

Treffer 1 - 1 | 1 Treffer