TAXN: Translate Align Extract Normalize, a Multilingual Extraction Tool for Clinical Texts

Neuraz, Antoine; Lerner, Ivan; Birot, Olivier; Arias, Camila; Han, Larry; Bonzel, Clara Lea; Cai, Tianxi; Huynh, Kim Tam; Coulet, Adrien

doi:10.3233/SHTI231045

loading subjects...

TAXN: Translate Align Extract Normalize, a Multilingual Extraction Tool for Clinical Texts

Authors

Antoine Neuraz, Ivan Lerner, Olivier Birot, Camila Arias, Larry Han, Clara Lea Bonzel, Tianxi Cai, Kim Tam Huynh, Adrien Coulet

Pages

649 - 653

DOI

10.3233/SHTI231045

Category

Research Article

Series

Studies in Health Technology and Informatics

Ebook

Volume 310: MEDINFO 2023 — The Future Is Accessible

Abstract

Several studies have shown that about 80% of the medical information in an electronic health record is only available through unstructured data. Resources such as medical terminologies in languages other than English are limited and restrain the NLP tools. We propose here to leverage English based resources in other languages using a combination of translation, word alignment, entity extraction and term normalization (TAXN). We implement this extraction pipeline in an open-source library called “medkit”. We demonstrate the interest of this approach through a specific use-case: enriching a phenotypic dictionary for post-acute sequelae in COVID-19 (PASC). TAXN proved to be efficient to propose new synonyms of UMLS terms using a corpus of 70 articles in French with 356 terms enriched with at least one validated new synonym. This study was based on freely available deep-learning models.

Contact

IOS Press Copyright 2024

Contact

IOS Press Copyright 2024

This website uses cookies

This website uses cookies