Advancing Italian Biomedical Information Extraction with Transformers-based Models: Methodological Insights and Multicenter Practical Application

Crema, Claudio; Buonocore, Tommaso Mario; Fostinelli, Silvia; Parimbelli, Enea; Verde, Federico; Fundarò, Cira; Manera, Marina; Ramusino, Matteo Cotta; Capelli, Marco; Costa, Alfredo; Binetti, Giuliano; Bellazzi, Riccardo; Redolfi, Alberto

doi:10.1016/j.jbi.2023.104557

Computer Science > Computation and Language

arXiv:2306.05323 (cs)

[Submitted on 8 Jun 2023 (v1), last revised 15 Jan 2024 (this version, v2)]

Title:Advancing Italian Biomedical Information Extraction with Transformers-based Models: Methodological Insights and Multicenter Practical Application

Authors:Claudio Crema, Tommaso Mario Buonocore, Silvia Fostinelli, Enea Parimbelli, Federico Verde, Cira Fundarò, Marina Manera, Matteo Cotta Ramusino, Marco Capelli, Alfredo Costa, Giuliano Binetti, Riccardo Bellazzi, Alberto Redolfi

View PDF

Abstract:The introduction of computerized medical records in hospitals has reduced burdensome activities like manual writing and information fetching. However, the data contained in medical records are still far underutilized, primarily because extracting data from unstructured textual medical records takes time and effort. Information Extraction, a subfield of Natural Language Processing, can help clinical practitioners overcome this limitation by using automated text-mining pipelines. In this work, we created the first Italian neuropsychiatric Named Entity Recognition dataset, PsyNIT, and used it to develop a Transformers-based model. Moreover, we collected and leveraged three external independent datasets to implement an effective multicenter model, with overall F1-score 84.77%, Precision 83.16%, Recall 86.44%. The lessons learned are: (i) the crucial role of a consistent annotation process and (ii) a fine-tuning strategy that combines classical methods with a "low-resource" approach. This allowed us to establish methodological guidelines that pave the way for Natural Language Processing studies in less-resourced languages.

Comments:	2 figures, 6 tables, Supplementary Notes included
Subjects:	Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
ACM classes:	I.2.7; J.3
Cite as:	arXiv:2306.05323 [cs.CL]
	(or arXiv:2306.05323v2 [cs.CL] for this version)
	https://doi.org/10.48550/arXiv.2306.05323
Journal reference:	Journal of Biomedical Informatics, Volume 148, 2023, 104557, ISSN 1532-0464
Related DOI:	https://doi.org/10.1016/j.jbi.2023.104557

Submission history

From: Tommaso Mario Buonocore [view email]
[v1] Thu, 8 Jun 2023 16:15:46 UTC (587 KB)
[v2] Mon, 15 Jan 2024 11:05:23 UTC (1,063 KB)

Computer Science > Computation and Language

Title:Advancing Italian Biomedical Information Extraction with Transformers-based Models: Methodological Insights and Multicenter Practical Application

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computation and Language

Title:Advancing Italian Biomedical Information Extraction with Transformers-based Models: Methodological Insights and Multicenter Practical Application

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators