Explainable Early Risk Identification of Parkinson’s Disease Using Multimodal Data Fusion With Machine Learning and Deep Learning Methods

Parkinson’s disease (PD) is a progressive neurodegenerative disorder whose early symptoms are often subtle and nonspecific. As global populations continue to age, the burden of PD is expected to increase substantially. Because large biomedical databases are primarily composed of general health surveys, clinical measurements, and genotype data rather than disease-specific diagnostic tests, identifying individuals at potential risk before clear clinical symptoms emerge remains an important challenge.

This study develops an explainable early risk prediction framework for Parkinson’s disease using data from the Taiwan Biobank. Genetic information, including single nucleotide polymorphisms (SNPs), polygenic risk scores (PRS), and principal components (PCs), is integrated with clinical information such as demographics, lifestyle factors, disease history, physical examinations, and biochemical measurements. Multiple machine learning and deep learning models are evaluated using both single-modal and multimodal datasets.

The study compares predictive performance from three perspectives: single-modal versus multimodal data, machine learning versus deep learning models, and different multimodal fusion strategies. The results show that clinical data provide a more stable single-modal information source, while integrating genetic and clinical features improves several important performance metrics, supporting the complementary value of the two data modalities. Among the machine learning models, Random Forest (RF) demonstrates strong overall discrimination and disease-sample identification ability, while TabM achieves the strongest overall performance among the deep learning models. In addition, Early Fusion, Intermediate Fusion, and Late Fusion show different strengths across evaluation metrics, indicating that the choice of fusion strategy affects the balance between overall classification performance and the ability to identify Parkinson’s disease samples.

Model interpretability is further examined using SHAP (SHapley Additive exPlanations). Important predictive features include SNPs, PRS, population principal components, clinical characteristics, lifestyle factors, and metabolic indicators. Several identified features are also consistent with previous Parkinson’s disease research, suggesting that multimodal learning can provide both predictive value and interpretable insights into potential disease risk factors.

Similar Posts