Juan M. Marquez-Romero, Hospital General de Zona No 2, Instituto Mexicano del Seguro Social (IMSS), OOAD Aguascalientes, Aguascalientes, Mexico
Svetlana V. Doubova, Unidad de Investigación Epidemiológica y Servicios de Salud Centro Médico Nacional Siglo XXI, IMSS, Ciudad de México, México
Flavio Cuellar-Roque, Hospital General de Zona No 2, Instituto Mexicano del Seguro Social (IMSS), OOAD Aguascalientes, Aguascalientes; Programa de Maestría y Doctorado en Ciencias Médicas Odontológicas y de la Salud, Universidad Nacional Autónoma de México (UNAM), Ciudad de México; México
Carlos A. Prado-Aguilar, Coordinación Medica Auxiliar de Investigación en Salud, IMMS, OOAD Aguascalientes, Aguascalientes, México
Alicia Alanis-Ocadiz, Unidad de Medicina Familiar No 8, IMSS, OOAD Aguascalientes, Aguascalientes, México
Jannett Padilla-López, Unidad de Medicina Familiar No 1, IMSS, OOAD Aguascalientes, Aguascalientes, México
Diana C. Navarro-Rodríguez, Hospital General de Zona No 1, IMSS, OOAD Aguascalientes, Aguascalientes, México
Carolina Quiñones-Villalobos, Programa de Maestría y Doctorado en Ciencias Médicas Odontológicas y de la Salud, Universidad Nacional Autónoma de México (UNAM), Ciudad de México, México
Ángel E. Muñoz-Zavala, Departamento de Estadística, Centro de Ciencias Básicas, Universidad Autónoma de Aguascalientes, Aguascalientes, Mexico
Rogelio Salinas-Gutiérrez, Departamento de Estadística, Centro de Ciencias Básicas, Universidad Autónoma de Aguascalientes, Aguascalientes, Mexico
Objective: The objective of the study is to determine the performance of three algorithms to classify patients with neurocognitive disorders during the pre-COVID-19 period. Methods: Data come from the electronic health records (EHRs) of patients hospitalized for COVID-19 aged > 18. EHRS entries before hospitalization were screened by an automated Python script that searched for evidence of neurocognitive testing during the pre-pandemic period. The neurocognitive disorder was determined based on the results of the neurocognitive testing. Machine learning (ML) classifiers (logistic regression [LR], k-nearest neighbors [k-NN], linear discriminant analysis [LDA]) were trained with the frequencies of appearance in the EHRs of a list of nine keywords obtained through a Delphi panel. The ten-fold cross-validated performance of classifiers was evaluated. Results: 3,674 EHRs were screened; 205 featured a pre-pandemic neurocognitive examination, and 127 were positive for neurocognitive disorder. The mean age was 81.8 ± 6.8 years (76% women). The mean number of keywords in the EHRs was 17.4. The most frequent keyword was a referral for suspicion of dementia. The three classifiers performed comparably in 12 out of 13 performance measures. Area under the curve (AUCs) were statistically different, with LR achieving the highest AUC (0.878), followed by k-NN (0.590) and LDA (0.524), p < 0.001. Conclusions: The three ML algorithms can accurately predict the presence of neurocognitive disorders in a sample of patients hospitalized for COVID-19 with consistent performance. The proposed approach can be used to determine the cognitive status of patients in various medical settings, including the pre-pandemic cognitive status of patients with suspected post-COVID-19 neurocognitive disorders.
Keywords: Neurocognitive disorders. COVID-19. Machine learning. Electronic health records.