Recognition and normalization of multilingual symptom entities using in-domain-adapted BERT models and classification layers
| dc.centro | E.T.S.I. Informática | es_ES |
| dc.contributor.author | Gallego Donoso, Fernando | |
| dc.contributor.author | Veredas-Navarro, Francisco Javier | |
| dc.date.accessioned | 2024-09-13T07:36:22Z | |
| dc.date.available | 2024-09-13T07:36:22Z | |
| dc.date.issued | 2024-08-28 | |
| dc.departamento | Lenguajes y Ciencias de la Computación | |
| dc.description.abstract | Due to the scarcity of available annotations in the biomedical domain, clinical natural language processing poses a substantial challenge, espe- cially when applied to low-resource languages. This paper presents our contributions for the detection and normalization of clinical entities corresponding to symptoms, signs, and findings present in multilingual clinical texts. For this purpose, the three subtasks proposed in the SympTEMIST shared task of the Biocreative VIII conference have been addressed. For Subtask 1—named entity recognition in a Spanish corpus—an approach focused on BERT-based model assemblies pretrained on a proprietary oncology corpus was followed. Subtasks 2 and 3 of SympTEMIST address named entity linking (NEL) in Spanish and multilingual corpora, respectively. Our approach to these subtasks followed a classification strategy that starts from a bi-encoder trained by contrastive learning, for which several SapBERT-like models are explored. To apply this NEL approach to different languages, we have trained these models by leveraging the knowledge base of domain-specific medical concepts in Spanish supplied by the organizers, which we have translated into the other languages of interest by using machine translation tools. | es_ES |
| dc.description.sponsorship | The authors acknowledge the support from the Ministerio de Ciencia e Innovación (MICINN) under project AEI/10.13039/501100011033. This work is also supported by the University of Malaga/CBUA funding for open access charge. | es_ES |
| dc.identifier.citation | Fernando Gallego, Francisco J Veredas, Recognition and normalization of multilingual symptom entities using in-domain-adapted BERT models and classification layers, Database, Volume 2024, 2024, baae087, https://doi.org/10.1093/database/baae087 | es_ES |
| dc.identifier.doi | 10.1093/database/baae087 | |
| dc.identifier.uri | https://hdl.handle.net/10630/32547 | |
| dc.language.iso | eng | es_ES |
| dc.rights | Atribución 4.0 Internacional | * |
| dc.rights.accessRights | open access | es_ES |
| dc.rights.uri | http://creativecommons.org/licenses/by/4.0/ | * |
| dc.subject | Informática - Salud | es_ES |
| dc.subject.other | Modelos BERT | es_ES |
| dc.subject.other | Registros de salud electrónicos | es_ES |
| dc.title | Recognition and normalization of multilingual symptom entities using in-domain-adapted BERT models and classification layers | es_ES |
| dc.type | journal article | es_ES |
| dc.type.hasVersion | VoR | es_ES |
| dspace.entity.type | Publication | |
| relation.isAuthorOfPublication | b8ab3a42-65ef-4349-9230-798e19f78426 | |
| relation.isAuthorOfPublication.latestForDiscovery | b8ab3a42-65ef-4349-9230-798e19f78426 |
Files
Original bundle
1 - 1 of 1
Loading...
- Name:
- Recognition_baae087.pdf
- Size:
- 10.62 MB
- Format:
- Adobe Portable Document Format
- Description:

