An Efficient Machine Learning and Deep Learning Framework for Early Prediction and Risk Assessment of Chronic Diseases
 
Pragathi1*, Dr. Neha Gupta2
1 Research Scholar, Department of Computer Science, Dr.K.N. Modi University Newai Tonk, Rajasthan, India
pragathi.vv@gmail.com
2 Associate Professor, Computer Science & Engineering, Dr.K.N. Modi University Newai Tonk, Rajasthan, India
Abstract: Timely identification of chronic illnesses is imperative in minimizing death rates, enhancing patient care, and facilitating early clinical treatment. This paper presents an effective machine learning and deep learning framework to early-stage predict the Chronic Kidney Disease (CKD) and chronic heart disease using structured clinical data available in the UCI Machine Learning Repository. The main aim of the study is to establish the perfect predictive models combined with the risk evaluation and personalized recommendation systems in the context of intelligent healthcare support.The suggested methodology will involve data preprocessing, missing value imputation, feature scaling, exploratory data analysis, feature selection and stratified data splitting. Several machine learning models were used to predict CKD including Naive Bayes, K-Nearest Neighbors, Decision Tree, Logistic Regression, Support Vector Machine models, and random forest models and Hybrid Ensemble models. Moreover, a more sophisticated CNN-UDRP deep learning network with embedding layers, convolutional neural network, and latent feature extraction via PCA was created. Deep ANN and Advanced DNN with LeakyReLU architectures were applied to predict chronic heart diseases with the help of batch normalization and dropout regularization.Experimental findings prove that the Random Forest model provided the highest performance of CKD prediction with the highest accuracy, precision, recall and F1-score of 100, whereas the Advanced DNN with LeakyReLU provided the highest accuracy of heart disease prediction at 99.93. The suggested framework also incorporates risk stratification and individual health recommendations, thus it can be used in scalable and intelligent healthcare decision-support systems.
Keywords: Chronic Disease Prediction, Deep Learning, Machine Learning, Chronic Kidney Disease and Heart Disease Prediction

1. INTRODUCTION

Chronic diseases like Chronic Kidney Disease (CKD) and cardiovascular disorders have emerged as significant healthcare issues of global concern following the rising morbidity, morbidity and mortality of these diseases. These diseases tend to be insidious and they are not diagnosed early enough before they cause serious health conditions, organ dysfunction and high healthcare expenses [1]. Unhealthy lifestyles, diabetes, high blood pressure and obesity, old age, and failure to diagnose early are among the factors that are driving the rapid increase in chronic diseases across the globe. Thus, prompt identification and precise forecast of chronic illnesses is crucial to good treatment, preventive medicine, and patient outcomes [2][3].
Conventional diagnosis of diseases largely relies on laboratory tests, physician knowledge and manual examination of the patient documentation [4], [5], [6], [7], [8], [9]. Despite the clinical importance of these methods, they can be costly and lengthy to implement and scaled to a large healthcare system [10]. The accelerating development of healthcare data produced by hospitals and electronic health records has presented the possibility to use machine learning and deep learning methods in the disease prediction and clinical decision support [11]. These smart methods can process high masses of medical information, discover the latent trends and aid in the prediction of diseases with accuracy and automation [12], [13].
Logistic Regression, Decision tree, K-Nearest neighbors, Support Vector machine, Naive Bayes, and random forest, are machine learning models that have demonstrated positive results in healthcare prediction [14][15], [16]. Nevertheless, traditional machine learning methods can be strongly dependent on hand feature engineering and can fail to capture the complex non linear relationships that can exist in heterogeneous clinical data [17], [18], [19], [20]. Deep learning methods address most of these constraints through automated learning of high-level representations of features of structured and unstructured healthcare data. Artificial Neural Networks (ANNs), Deep Neural Networks (DNNs), and Convolutional Neural Networks (CNNs) are architectures that have shown good performance in enhancing predictive accuracy and processing complex medical data [21], [22], [23], [24].
Nevertheless, much of the current research concentrates on a specific disease, has no solid preprocessing and feature optimization procedures, or does not offer individualized healthcare recommendations upon prediction. Also, processing of mixed clinical data which includes numerical, categorical and text-like data is a difficult issue in smart healthcare [25], [26]. To overcome these shortcomings, this paper presents an effective model to predict chronic diseases early on based on machine learning and deep learning algorithms, namely Chronic Kidney Disease and chronic heart disease [27], [28], [29], [30].
The framework suggested combines the benefits of cutting-edge preprocessing, feature selection, dimensionality reduction, machine learning classifiers, and deep learning architectures to enhance the accuracy of predictions and the ability to generalize. To predict CKD, several machine learning models are deployed and evaluated in addition to an advanced CNN-based Unified Disease Risk Prediction (CNN-UDRP) model that can learn latent semantic patterns on clinical representations. In prediction of heart disease, hybrid deep learning models using ANN and state-of-the-art DNN with LeakyReLU activation are created to enhance the feature learning and classification stability.
This study is novel because it integrates structured clinical analysis, latent feature extraction (Principal Component Analysis (PCA)) with clinical text representation and deep learning into a single healthcare system. The proposed approach, unlike traditional prediction systems, also incorporates the personalized risk assessment and recommendation system that classifies patients into risk levels and offers disease-specific health care advice. This converts the framework into an intelligent clinical decision-support model as opposed to a mere prediction system.
The key findings of the study are the creation of predictive models of chronic diseases based on machine learning and deep learning, the implementation of advanced CNN and DNN architecture, combining the risk stratification and recommendation systems, and improving predictive healthcare analytics to support large-scale real-world medical tasks. The suggested framework will help in the facilitation of early diagnosis, mitigation of disease progression risks, enhancing preventive healthcare strategies, and help advance intelligent healthcare systems.

2. LITERATURE REVIEW

Saxena 2026 et al. presented a new hybrid machine learning model -XGBoost- LSTM-Attention (XLA) - that incorporates gradient-boosted feature selection with long short- term memory (LSTM) networks and a temporal attention mechanism to predict CKD progression between Stage 3 and Stages 4/5 or ESRD using longitudinal claims-based features. Primary analysis: a cross-sectional validation with real NHANES 2015 2018 data (n=701 CKD Stage 3 adults) by predicting significant proteinuria (UACR ≥30 mg/g) based on clinical features without UACR. Additional analysis: an NHANES-calibrated longitudinal cohort (n=8,412) simulated to have quarterly measurements showed XLA performance in real-world longitudinal data conditions. All models were tested on the basis of 5-fold stratified cross-validation. Findings: The XLA framework in the primary NHANES cross-sectional analysis yielded AUC- ROC of 0.684 (95% CI: 0.6410.727), and all of the models were similar (AUC 0.684 -0.710), which confirms that cross-sectional clinical characteristics alone do not give much signal about the prediction of proteinuria and the importance of UACR measurements. In longitudinal supplementary analysis, XLA demonstrated AUC-ROC of 0.994 compared to 0.939 at the best cross- sectional baseline of +5.5% indicating that features of temporal trajectories when longitudinal data are present, especially the slope of eGFR and the trends of RAAS adherence, provide considerable incremental predictive value. Conclusion: XLA framework shows significant benefits compared to the conventional methods when used with longitudinal claims data. Cross-sectional results emphasize the inability of alternative direct UACR measurement to substitute the direct measurement of UACR in risk stratification of CKD. Collectively, these outcomes offer practical evidence to both the shortcomings of the traditional prediction and the potential of trajectory-oriented strategies of value-based care programs dealing with large CKD cohorts [31].
Al Ghamdi & Alyahyan, 2026 et al. it presents three major innovations: (1) a hierarchical attention model that learns the complex inter-dependencies between clinical parameters, (2) an adaptive feature fusion model that integrates patterns learned with transformers with gradient-boosting decision boundaries, and (3) a confidence-aware ensemble strategy that quantifies uncertainty in clinical decision support. DeepCKD-Net has an accuracy of 98.7 and AUC of 0.993 which is 4.2 per cent higher than state-of-the-art methods with an inference time of 16.8 ms, which can be deployed in real-time clinical settings. Integrated SHAP analysis offers the possibility of interpretable predictions and the best predictive biomarkers are serum creatinine (SHAP value: 0.342) and blood urea (0.287), which are consistent with the clinical knowledge. The framework has shown good performance in real clinical settings and is accurate with more than 90 percent and 20 percent of missing data. We have contributed to the development of AI-based nephrology diagnostics through the provision of a deployable, interpretable, and clinically validated early CKD detection solution [32]
Miller 2026 et al. It was proposed by introduced a Multi-Perspective Machine Learning system (MPML), as a remedy to this, using known base classifiers with a hierarchical perspective-based framework and interpretability pipeline. MPML allows instance and global interpretability by being able to group the features in meaningful subsets, or perspectives. On evaluatory measures, MPML is predicting better than classic ensemble predictive algorithms like Bagging, Boosting and Random Forest, as well as, it is transparent. With the heart disease data, MPML not only improves the quality of prediction, but also gives detailed and understandable explanations of individual predictions, which is a feasible, robust and ethical approach to use machine learning in healthcare and balance the high-performance accuracy with clinical credibility [33].
Bhatta 2026 et al. identified six most popular ML models, such as Logistic Regression (LR), Support Vector Machine (SVM), Random Forest (RF), Gradient Boosting (GB), K-Nearest Neighbors (KNN) and eXtreme Gradient Boosting (XGB) models on a publicly available Kaggles heart disease dataset to compare them on a standard preprocessing dataset. Accuracy, precision, recall and F1-score were used to measure the performance. Random Forest and XGBoost performed better with an F1-score of 98.52, an accuracy of 98.53, and a perfect rate (100%), and recall of over 97, and Gradient Boosting had an accuracy of 93.17. These findings suggest that it is possible to use ML models, in particular RF and XGB, to support the clinical decision-making process, improve cardiovascular Diagnosis and develop awareness of heart health [34]
Ashish Kumar 2025 et al. The statistics of 1,190 cases of 11 common characteristics in five official sites, specifically, Cleveland, Hungarian, Switzerland, Long Beach, Virginia and Statlog form the basis of the paper as the most comprehensive data on the study of coronary artery disease (CAD). The null values were verified with the help of exploratory data analysis and split the dataset into 80/20 train/test and traditional scaling to possess the same features. Eight machine learning models were run and they included the Logistic Regression, Decision Tree, Random Forest, Support Vector machine, K- nearest neighbours, Gradient Boosting, AdaBoost and XGBoost. The best test accuracy (0.966), precision (0.967) and recall (0.966) were achieved by XGBoost that is a supported predictor via grid search and 5-fold cross-validation, and can be used to obtain reliable data on CAD prediction and early diagnosis of the disease in clinical practice [35].
Table 1: Literature Summary
Author / Year
Method Used
Key Results
Research Gap Identified
Dey & Kamal, 2026 (Dey & Kamal, 2026)
Proposed a machine learning framework for CKD prediction using feature significance analysis through heatmap correlation, information gain, and standard deviation. Logistic Regression (LR), Support Vector Machine (SVM), and Multi-Layer Perceptron (MLP) classifiers were implemented with selective feature elimination.
Average feature scores achieved 100% classification accuracy. Logistic Regression obtained 100% accuracy after removing nine non-critical features. Maximum-value-based evaluation achieved 96.75% accuracy with reduced computational complexity and efficient feature optimization.
The study mainly focused on structured numerical data and traditional machine learning models. Deep learning-based semantic feature extraction, clinical text representation, risk assessment, and personalized recommendation systems were not explored.
Khafaga et al., 2025 (Khafaga et al., 2025)
Developed a deep neural network using a hybrid optimization framework combining Waterwheel Plant Algorithm (WWPA) and Grey Wolf Optimization (GWO). Preprocessing included imputation, normalization, and synthetic oversampling on the UCI CKD dataset.
The optimized MLP significantly improved prediction performance with reduced error metrics and achieved higher prediction accuracy and robustness compared to PSO, GA, and WOA optimization techniques. Statistical validation confirmed model reliability and reduced computational time.
The framework focused primarily on optimization of prediction accuracy and lacked integration of explainable AI, personalized healthcare recommendations, and real-world clinical decision support systems. Multi-disease prediction capability was also not addressed.
Ahmed/2024[36]
Ensemble and hybrid machine learning applied on cardiac datasets.
Improved precision and reliability in heart disease prediction outcomes.
Need for more robust models combining multiple predictive algorithms.
Manikandan/2024[37]
Logistic Regression, Decision Tree, SVM with Boruta feature selection
Boruta improved performance; Logistic Regression achieved 88.52% accuracy.
Need for exploring additional feature selection methods for higher accuracy.
Lingadally /2024[38]
Random Forest, MLP, XGBoost, LightGBM, stacking ensemble, logistic regression
Highest accuracy 95.8%, outperforming other machine learning techniques
Further improvement, broader validation, and real-world clinical implementation needed
Kellen /2023[39]
Random Forest, feature selection, correlation, data preprocessing, decision trees, ensemble
Achieved 99% accuracy, outperforming K-NN, SVM, Logistic Regression
Validation on larger, diverse datasets and real-world clinical implementation needed
 
Loveleen /2023[40]
Convolutional Neural Network, ECG signals, demographic, clinical data, deep learning
Achieved over 90% accuracy, outperforming classical AI and alternative models
Further validation on larger datasets and integration into real-world systems
 

3. METHODOLOGY

The proposed methodology can be divided into two large sections to develop and test an effective deep learning framework to predict chronic diseases early systematically.
6. Evaluation using accuracy, Precision, Recall, F1 score, ROC
 
Figure 1: Proposed Flowchart

Part 1: Chronic Kidney Disease (Ckd) Prediction Framework

3.1 Data Collection

Data in the present study was taken to the UCI Machine Learning Repository, namely, the Chronic Kidney Disease (CKD) dataset, which is a standardized benchmark dataset used in medical prediction studies. The data set was initially obtained as clinical records of patients in the course of regular medical check-ups and laboratory tests, which has guaranteed its real-life applicability and clinical soundness.
It comprises of 400 patient cases having 25 attributes among which 24 predictive variables and one binary outcome variable denoting the presence or the absence of chronic kidney disease. The features collected are the holistic product of demographic, clinical observations, and biochemical test outcomes. These are: age, blood pressure, specific gravity, albumin, sugar, red and pus cell factors, blood glucose random, blood urea, serum creatinine, sodium, potassium, hemoglobin, packed cell volume, white and red blood cell count, and indicators of comorbidity like hypertension, diabetes mellitus, and coronary artery disease. There are also both numeric (integer and continuous) and categorical/binary variables as medical data can be very heterogeneous. The target attribute, denoted class, will classify patients as CKD and non-CKD. This multi-dimensional and clinically rich data allows developing strong deep learning models to predict chronic diseases early and correctly, especially chronic kidney disease.

3.2 Data Preprocessing

Data preprocessing plays an important role in an effective and stable deep learning system to predict chronic diseases early because medical data may include missing values, different formats, mixed data types, and feature scale differences, which can adversely affect model results. In this research, a multi-stage preprocessing pipeline was implemented in a structured fashion to the Chronic Kidney Disease (CKD) dataset to improve the quality of the data and its consistency to be used in training the model. To begin with, the target variable was first filtered by removing redundant spaces and transformed into binary, with CKD being 1 and non-CKD being 0 to simplify the supervised learning process. The analysis of missing values followed, with imputation plans determined depending on the feature type; numerical values were imputed with median values to cope with outliers and categorical values were imputed with the most common value to maintain the most common clinical trends. Label encoding was also used to convert categorical variables into numerical format so that they can be compatible with deep learning models. In order to deal with variations in features of different magnitudes, the data was normalized by zero mean and unit variance to guarantee equal contribution of the all variables in the training process and enhance the stability of converging. The clean dataset that followed preprocessing was structured and consistent to the full extent to develop the model. Moreover, feature selection with Mutual Information and SelectKBest was used to select the most relevant predictors and to reduce dimensionality and enhance the model efficiency, interpretability, and generalization to enable accurate predicting of CKD.

3.3 Exploratory Data Analysis

Exploitative Data Analysis shows significant trends about CKD data such as high correlation of hypertension and diabetes with CKD, elevated serum creatinine level among CKD patients, unbalanced classes between CKD and non-CKD patients, and meaningful associations between various clinical and laboratory variables that could be effective in the feature selection.
Figure 2: Hypertension and diabetes prevalence in CKD
CKD data reveals a mediocre hypertension correlation and greater diabetes connection with type 1 cases predominating over type. 2.
Figure 3: Serum Creatinine by Disease Class
The creatinine rates of serum are more varied and elevated in CKD class, which indicates poor kidney functionality in comparison with the non-CKD class.
Figure 4: CKD vs Non-CKD Distribution
The number of patients in CKD class is higher than in non-CKD, pointing to the imbalance of the dataset and the increased prevalence of chronic kidney disease.
Figure 5: Correlation Heatmap of CKD Features
Heatmap shows that lab values and clinical features have strong relationships, which will be used to select features to predict and diagnose CKD.

3.4 Data Splitting

Splitting the data is a critical aspect of the creation of a robust deep learning-based prediction model because it provides objective assessment of unknown data. In this research, the processed data set was separated into training and testing data sets in a ratio of 80:20. Training set was employed to learn the model and optimize the parameters, and the testing set was utilized as an independent performance test. To guarantee reproducibility of results, a fixed random state (42) was used. Besides, stratified sampling was applied to preserve the initial distributions of CKD and non-CKD cases in both subsets to resolve the problem of imbalance in classes. This method guarantees equitable training, less bias, and offers a realistic evaluation of the generalization capacity of the model in clinical prediction activities.

3.5 Model Implementation

3.5.1 Machine Learning Model Implementation

In order to access the efficiency of various machine learning methods in early chronic kidney disease prediction, various classification models were developed and compared on the basis of the preprocessed and scaled dataset. These models were chosen according to their effectiveness in the medical diagnosis task, capacity to deal with nonlinear relationships, and interpretability.
  1. Naïve Bayes Classifier
Gaussian Naive Bayes (GNB) model was used as a probabilistic baseline classifier. It makes the assumption of conditional independence between features and models continuous attributes on a Gaussian distribution. Naive Bayes, in spite of its simplicity, has been shown to be computationally efficient and works well on high-dimensional medical data and is therefore appropriate in initial prediction of CKD.
The conditional probability in Gaussian Naïve Bayes is:
P(Ck∣x)=P(Ck)i=1nP(xi∣Ck)P(x) (1)
Gaussian distribution for continuous features:
P(xi∣Ck)=12πσk2exp⁡xi-μk22σk2 (2)
2. K-Nearest Neighbors (KNN)
K-Nearest Neighbors classifier was utilized to obtain instance-based similarities among patients. The model has a size of k = 5, which predicts the class of a test instance as the majority vote of the nearest neighbors of the test instance in the feature space. KNN can be used to determine local patterns in clinical data but is susceptible to scale of features and this was considered during preprocessing.
Euclidean distance measure:
d(x,y)=i=1n(xi-yi)2 (3)
3. Decision Tree Classifier
An entropy criterion was used with a maximum depth of 6 to create a Decision Tree model. This limitation aids in overfitting control and maintain interpretability. Decision Trees have been especially useful in healthcare applications because of their rule-based decision paths that are easy to follow and generate clear decisions that are consistent with clinical reasoning.
Entropy equation:
Entropy(S)=-i=1cpilog⁡2(pi) (4)
Information gain:
IG(S,A)=Entropy(S)-v∈Values(A)∣Sv∣∣S∣Entropy(Sv) (5)
4. Logistic Regression
A linear classification model was applied, using Logistic Regression to predict the likelihood of developing CKD. The model has a maximum of 1000 iterations, which ensures convergence and is very interpretable and robust, particularly with linearly separable patterns.
Sigmoid function:
P(y=1∣x)=11+e-(wTx+b) (6)
5. Support Vector Machine (SVM)
The Support Vector Machine (with the Radial Basis Function (RBF) kernel) was used to predict the existence of complex nonlinear correlations in the dataset. The probability estimation option was turned on to permit probabilistic outputs, necessary to ensemble learning and clinical risk assessment.
Hyperplane equation:
f(x)=wTx+b (7)
RBF kernel:
K(x,x')=exp⁡(-γ∣∣x-x'∣∣2) (8)
6. Random Forest Classifier
Random Forest model has been created to be an ensemble of 100 decision trees with a maximum depth of 8. The model combines the predictions of many trees to minimize variance, enhance generalization, and effectively model nonlinear interactions of features exhibited in medical data.
Ensemble prediction:
y=1Bb=1Bfb(x) (9)
7. Hybrid Ensemble Model
In order to improve predictive performance, soft voting ensemble model was built by using Naivete Bayes, Logistic Regression, and Random Forest classifiers. VotingClassifier is used to combine predicted class probabilities of individual models and result in a more stable and accurate prediction. This combination of methods capitalizes on the advantages of probabilistic and linear learners as well as ensemble-based methods, which makes the combination suitable to make dependable early forecasts of chronic kidney disease.

3.5.2 Advanced Deep Learning Framework

In order to thoroughly assess the efficacy of various machine learning and deep learning models to help predict the onset of chronic kidney disease (CKD) in its initial stages, the proposed research adopted a combination modeling approach, combining traditional machine learning classifiers, dimensionality reduction, and a state-of-the-art deep learning architecture. It was proposed that the suggested framework would be able to include both structured clinical measurements and latent representations based on patient-specific information and, thus, enhance predictive robustness and clinical relevance.
  1. Latent Feature Extraction Using PCA
Following the preprocessing and scaling of the features, the Principal Component Analysis (PCA) was used to the numerical set of features to identify latent representations and reduce the dimensions. PCA was set to retain 10 principal components, which retained most of the variance without the redundancy and multicollinearity in clinical attributes. These latent representations gave a concise and noise-free embodiment of patient health indicators, which supported the efficient learning of conventional machine learning classifiers, as well as minimized computational complexity.
Principal Component Analysis (PCA)
Covariance matrix:
C=1n-1XTX (10)
Principal component transformation:
Z=XW (11)
2) Clinical Text Representation
Besides numerical data, a lightweight clinical text representation was built that approximated unstructured clinical narratives. The most important attributes of the patients like age, blood pressure, hypertension status, and diabetes status were combined in a text format, per patient record. This strategy allowed simulating contextual relationships among clinical indicators in a closer way to the actual electronic health records.
A Tokenizer was used to tokenize the generated clinical text with a vocabulary size of 5,000 words and an out-of-vocabulary token to treat unseen words. The sequences were finally tokenized and padded to an equal length of 50 tokens, guaranteeing equal dimensions of the deep learning model inputs.
 
3) Advanced CNN-UDRP Deep Learning Model
A more complex CNN-based Unified Disease Risk Prediction (CNN-UDRP) model was suggested to capture more patterns and complex dependencies in clinical text representations. The architecture starts with an explicit input layer then an Embedding layer (input dimension = 5,000, output dimension = 128) that learns dense semantic representations of clinical tokens.
Convolution Operation (CNN)
1D convolution equation:
hi=fj=0k-1wjxi+jb (12)
Two 1D convolutional layers with a filter size of 128 and 64 and a kernel size of 5 and 3, respectively, were used to process the embedded sequences. To stabilize the learning process and prevent overfitting, batch normalization and dropout layers were used. The most salient features in the entire sequence were extracted using a Global Max Pooling layer, guaranteeing the translation invariance and dimensionality reduction.
The convolutional feature maps were fed into fully connected layers with 128 and 64 neurons with dropout regularization. The last output layer used a sigmoid activation function to estimate the likelihood of CKD presence. The Adam optimizer plus a learning rate of 0.0005 and binary cross-entropy loss were used to compile the model. The accuracy, precision, recall and AUC measures were used to assess performance and thus give a global view of diagnostic performance.
Embedding Layer
Word embedding representation:
E(wi)∈Rd (13)
4) Model Complexity and Training Efficiency
The CNN- UDRP model was comprised of about 764, 097 parameters, most of which were trainable. Such a balanced complexity enabled the model to acquire rich representations with minimal computational overhead, which is why it is feasible to apply to actual healthcare problems.
Table 2: Hyperparameter Details of Implemented Models
Model
Hyperparameters
Selected Values
Gaussian Naïve Bayes (GNB)
var_smoothing
1e−9 (default)
K-Nearest Neighbors (KNN)
n_neighbors
5
 
Weights
Uniform
Decision Tree (DT)
criterion
Entropy
 
max_depth
6
 
min_samples_split
Default (2)
Logistic Regression (LR)
max_iter
1000
 
Regularization
L2
 
Solver
lbfgs
Support Vector Machine (SVM)
kernel
RBF
 
C
Default (1.0)
 
gamma
Scale
 
probability
True
Random Forest (RF)
n_estimators
100
 
max_depth
8
 
random_state
42
Hybrid Voting Classifier
Voting type
Soft voting
 
Base estimators
NB, LR, RF
CNN-UDRP Model
Embedding dimension
128
 
Conv1D filters
128, 64
 
Kernel sizes
5, 3
 
Dropout rates
0.3, 0.4
 
Optimizer
Adam
 
Learning rate
0.0005
 
Loss function
Binary cross-entropy
 
Evaluation metrics
Accuracy, Precision, Recall, AUC
  1. Soft Voting Ensemble
Soft voting probability aggregation:
y=arg⁡max⁡ij=1mwjPij (14)

Part 2: Chronic Heart Disease Prediction Framework

3.6 Data Collection

The current paper is based on a secondary clinical data acquired in UCI Machine Learning Repository, which is a reputable and authoritative source of benchmark healthcare data used in predictive modeling studies. In particular, the Heart Disease dataset (ID: 45) was utilized, because it has clinically relevant attributes that can be used to predict chronic cardiovascular disease early. The data set includes records of patients, each of which is a single subject, that experienced diagnostic testing related to heart disease.
The obtained data have 14 attributes that include demographic, clinical, and physiological parameters. The major variables are age and sex, medical factors like the type of chest pain, blood pressure during rest, serum cholesterol, fasting blood sugar, electrocardiographic results during exercise, peak exercise rate, angina during physical activity, ST depression (oldpeak), the slope of the peak exercise ST segment, excessive vessels filled with color after fluoroscopy, and thalassemia. The outcome variable represents whether or not there exists severity of heart disease and this was later converted into binary classification label to be used in predictive modeling.The privacy and ethical considerations were promoted by anonymizing all patient records. No further ethical approval was necessary as the dataset is publicly accessible and does not include any personal identifiable information.

3.7 Data Preprocessing

The preprocessing of heart disease data consisted of multiple crucial steps to prepare the data with high quality to be fed into deep learning models. The multi-class target variable (num 0-4) was first changed into binary format, with 0 meaning no heart disease and 1-4 meaning presence of the disease and then the column was removed to prevent redundancy. Numerical attributes that had missing values were analyzed and imputed with median values to lessen the impact of outliers. StandardScaler was used to perform feature scaling to normalize the variables, including age, cholesterol, and blood pressure, to be of zero mean and unit variance. Class imbalance was also studied by calculating class weights to eliminate preference to the majority classes in the training process. Lastly, the data was cleaned, normalized, and ready to be effectively predicted using deep learning in heart diseases.

3.8 Exploratory Data Analysis

Figure 6: Heart Disease Distribution
Dataset shows 54.1% without heart disease and 45.9% with disease, indicating nearly balanced distribution across studied population.
Figure 7: Age vs Heart Disease
Individuals with heart disease tend to be older, median age higher than non-disease group, highlighting age as key risk factor.
Figure 8: Feature Correlation Heatmap
Heatmap reveals strong correlations between clinical variables and heart disease, guiding feature selection for predictive modeling and diagnostic analysis.
Figure 9: Cholesterol vs Max Heart Rate
Figure 10: Sex, Chest Pain & Exercise Angina
Males show higher disease cases, chest pain types vary in association, and exercise-induced angina strongly correlates with heart disease.

3.9 Data Splitting

Stratified sampling was used to split the preprocessed dataset of heart disease into 80 percent training and 20 percent test to maintain the distribution of classes. This strategy guaranteed fair model testing, minimized the effects of class imbalance, enhanced the generalization ability, and reproducibility as it used a fixed random seed during data partitioning.

3.10 Hybrid Deep Learning Model Implementation

Two hybrid deep learning models were used to forecast the early development of chronic heart disease. These architectures exploit multilayer artificial neural networks with powerful regularization and activation schemes to achieve maximum predictive performance.
  1. Model 1 Deep Artificial Neural Network (ANN)
The initial architecture, Model 1, is a deep ANN that is implemented with several fully connected layers. The input layer has 128 neurons, using ReLU (Rectified Linear Unit) activation function, which enables the model to learn non-linear dependencies in the data without the vanishing gradient issue. ReLU is a popular deep learning operator because of ease of use and the ability to hasten convergence in training.
The model has two hidden layers with 64 and 32 neurons respectively after the input layer. The hidden layers also have a batch normalization that normalizes the inputs of every single layer, which enhances the stability as well as makes the learning process faster. To minimize overfitting, the dropout regularization is used with different probabilities (0.2-0.4) whereby in training, a certain portion of neurons are randomly disconnected so that the model does not depend on any one feature too much.
The output layer has a single neuron whose activation value is the sigmoid type, which is suited to a binary classification task because it gives a probability value of 0-1 about the presence of heart disease. The model is also based on the Adam optimizer with a learning rate of 0.001, which is selected due to its ability to adapt to learning and effective processing of sparse gradients. The loss function is binary cross-entropy, which quantifies the difference between predicted and actual class labels.
Model 1 has around 13,000 trainable parameters, which enables it to acquire complex patterns in the data set with a manageable computational footprint. Assessment of this architecture shows that it can obtain the necessary demographic, clinical, and physiological associations that lead to the risk of heart disease.
2. Model 2 Advanced Deep Neural Network (DNN) with LeakyReLU
The second architecture, the Model 2, is an improved DNN that is more effective in feature extraction and addressing some of the drawbacks of standard ANNs. In particular, this model uses LeakyReLU as the activation function in hidden layers. LeakyReLU, in contrast to a standard ReLU, enables a non-zero and small gradient on negative inputs, which can alleviate the dying ReLU problem, as neurons become inactive and do not contribute to learning anymore.
Model 2 has more input layer that has 256 neurons, then 128 and 64 neurons in the hidden layers, respectively. The batch normalization is implemented after every hidden layer to stabilize the learning process and enhance convergence. In this architecture, dropout rates are a bit higher (0.3 -0.5) to further minimize overfitting, especially because of the higher model complexity and number of parameters.
The output layer, just like that of Model 1, has a single neuron with sigmoid activation in binary classification. Adam is the optimizer, and the learning rate is reduced to 0.0005, which indicates that the model is more complex and requires more precise changes during the training process. Binary cross-entropy remains the loss function, and the model performance is based on accuracy.
The sophisticated DNN has more trainable parameters than Model 1 and can learn more complicated interactions between patient characteristics age, type of chest pain, resting blood pressure, cholesterol levels, and other cardiovascular measurements. LeakyReLU, used together with batch normalization and dropout, is robustly trained, generalizes better, and is able to address class imbalance.
Table 3: Hyperparameter details of proposed Models
Hyperparameter
Deep ANN
Advanced DNN with LeakyReLU
Input Layer Neurons
128
256
Hidden Layers
2 hidden layers: 64, 32 neurons
2 hidden layers: 128, 64 neurons
Activation Function
ReLU
LeakyReLU α=0.1
Dropout Rate
0.2–0.4
0.3–0.5
Optimizer
Adam
Adam
Learning Rate
0.001
0.0005
Loss Function
Binary Cross-Entropy
Binary Cross-Entropy
Output Layer
1 neuron, Sigmoid
1 neuron, Sigmoid
 
  1. ReLU Activation
f(x)=max⁡(0,x) (15)
2) LeakyReLU Activation
f(x)=x,x>0αx,x≤0 (16)
3) Sigmoid Activation
σ(x)=11+e-x (17 )
4) Binary Cross-Entropy Loss
L=-1Ni=1N[yilog⁡(yi)+(1-yi)log⁡(1-yi)] (18)
5) Adam Optimizer Update
Parameter update:
θt+1=θt-ηmtvt+ϵ (19)

4. RESULT & DISCUSSION

This part of the paper introduces and analyzes the performance of machine learning and deep learning models to predict chronic kidney and heart disease based on the performance metrics: accuracy, precision, recall, F1-score, and AUC-ROC. It also outlines the suggested risk assessment and customized recommendation frameworks to aid in early diagnosis, preventive care, and clinical decision-making.
  1. Accuracy: Measures overall correctly predicted instances among total predictions made by model.
Accuracy=TP+TNTP+TN+FP+FN (20)
2) Precision: Measures proportion of correctly predicted positive cases among predicted positives.
Precision=TPTP+FP (21)
3) Recall: Measures ability to correctly identify actual positive disease cases accurately.
Recall=TPTP+FN (22)
4) F1-Score: Harmonic mean balancing both precision and recall for classification performance.
F1=2×Precision×RecallPrecision+Recall (23)
5) AUC-ROC: Evaluates classification capability using true positive and false positive rates.
ROC relation:
TPR=TPTP+FN,FPR=FPFP+TN (24)

4.1 Performance Evaluation and Risk Assessment of CKD Prediction Models

Table 4: Performance Comparison of Machine Learning Models for CKD Prediction
Model
Accuracy
Precision
Recall
F1-Score
Naïve Bayes
0.9625
1.00
0.94
0.9691
K-Nearest Neighbors (KNN)
0.9875
1.00
0.98
0.9899
Decision Tree
0.9750
1.00
0.96
0.9796
Logistic Regression
0.9750
1.00
0.96
0.9796
Support Vector Machine (SVM)
0.9875
1.00
0.98
0.9899
Random Forest
1.0000
1.00
1.00
1.0000
Hybrid Ensemble Model
0.9750
1.00
0.96
0.9796
 
Table 4.1 compares the performance of various machine learning models on chronic kidney disease prediction in terms of accuracy, precision, recall, and F1-score. The Random Forest classifier performed best with the highest accuracy, precision, recall and F1-score rate of 1.0000 meaning it is outstanding, as it was able to identify correctly the number of CKD and non-CKD cases. K-Nearest Neighbors (KNN) and Support Vector Machine (SVM) also yielded very promising results with 98.75 percent accuracy and high recall scores meaning that they had good classification. Decision Tree, Logistic Regression and the Hybrid Ensemble model had a consistent results of 97.50% accuracy and equal precision-recall. Naïve Bayes had a slightly lower recall than other models, but the overall performance was high. The findings suggest that ensemble and nonlinear classifiers are very effective in the prediction of CKD by structured clinical data.
Figure 11: Performance Comparison of Machine Learning Models for CKD Prediction Graph
Table 5: Advanced Performance Evaluation of CNN-UDRP Deep Learning Model for CKD Prediction
Model Architecture
Accuracy
Precision
Recall
F1-Score
CNN-UDRP Model
0.9400
0.9669
0.9360
0.9512
 
Table 4.2 discusses the results of the research of the suggested CNN-UDRP deep learning network to predict CKD. The model obtained an accuracy of 94.00 as well as high levels of precision, recall, and F1-score, indicating its potential to detect cases of chronic kidney disease with a high level of accuracy and reducing classification mistakes. The high value of precision implies fewer false-positive predictions, whereas the high value of recall represents the model sensitivity to the ability to identify the real CKD patients. The F1-score of balance also demonstrates consistent and predictable prediction even when the clinical data is unbalanced. Even though the traditional machine learning models have had slightly better accuracy, the CNN-UDRP model has the added benefit of learning latent semantic and contextual patterns on clinical representations and therefore is applicable in scalable and real world healthcare applications, where there is complex medical data.
Figure 12: Performance Evaluation of CNN-UDRP Deep Learning Model for CKD Prediction
CNN-UDRP model exhibits a high level of precision and recall, which is an indication of its ability to reduce false positives whilst being very sensitive to cases of chronic kidney disease. The high F1-score indicates consistent results with uneven clinical data. Even though the ensemble machine learning models scored slightly better, CNN-based architecture is significantly better at learning latent semantic representations on clinical narratives, which is especially useful with scalable, real-world healthcare systems in which unstructured patient data is common.
Figure 13: Confusion Matrix of CNN-UDRP Deep Learning Model
Figure 14: Performance Graphs

4.1.1 Effective Risk Assessment and Personalized Recommendations

In the present research, the efficient risk assessment system and the individual recommendation system were created to help the discovery of chronic kidney disease (CKD) in the early stages with the help of machine learning and deep learning algorithms. Based on model forecasts, the likelihood of a patient developing CKD is computed and categorized as Low, Medium or High Risk. This stratification would guarantee that patients and clinicians were able to quickly determine the severity of the possible kidney dysfunction and prioritize interventions based on that severity. The predicted probabilities and the associated risk levels are then structured into a formatted DataFrame to make them easily interpretable to allow their downstream clinical decision-making processes.
On the basis of risk stratification, a disease-specific recommendation system was built in. In patients who have been predicted to have CKD, the choice varies between healthy lifestyle and periodic check-ups in cases that are low-risk and urgent medical visit and frequent kidney functional evaluation in high-risk cases. Prevention is given to non-CKD patients, such as periodic examinations and the optimization of lifestyle. The recommendations are dynamically produced based on a function that translates the predicted disease and level of risk to actionable advice.
Figure 15: Final output
The last output integrates the disease prediction, risk level, and personalized recommendations into a unified comprehensive report. Combined with predictive modeling and customized clinical guidance, this framework facilitates early detection, prevention, and decision-making, improving patient outcomes and facilitating proactive healthcare management.

4.2 Performance Evaluation and Personalized Recommendation System for Chronic Heart Disease Prediction

Table 6: Performance comparison of the proposed deep learning models for early prediction of chronic heart disease.
Model Name
Accuracy
Loss
Precision
Recall
F1-Score
Deep Artificial Neural Network
0.9903
0.0342
1.0000
1.0000
1.0000
Advanced Deep Neural Network with LeakyReLU
0.9993
0.0091
1.0000
1.0000
1.0000
 
Table 4.3 indicates that both of the proposed deep learning models have performed remarkably well in prediction of chronic heart diseases. The LeakyReLU version of the Advanced Deep Neural Network was slightly more effective than the regular ANN since it got a better result of accuracy and reduced loss. Both models were able to yield a perfect precision, recall and F1-score which shows a very reliable and accurate classification ability.
Figure 16: Deep learning models

4.2.1 Effective Risk Assessment and Personalized Recommendations

The proposed research incorporates the personalized recommendation system into the predictive model of chronic heart disease that offers actionable recommendations based on the personal risk level. The deep learning model estimates the probability of heart disease after the patient data has been preprocessed and the features were scaled with the trained StandardScaler. The model provides a binary classification of the presence or absence of heart disease-and a probability score of the risk.
The system categorizes risk as low, moderate or high, depending on the prediction and the probability determined. In the case of patients who are not at high risk, the system recommends some general prevention steps, such as a healthy diet, physical exercise, avoidance of smoking and alcohol, and regular health check-ups. In case of moderate-high risk patients, a more immediate advice is given, including consultation with a cardiologist, controlling blood pressure and cholesterol levels, taking medications strictly as prescribed, following a low-fat and low-sodium diet, engaging in doctor-prescribed physical activity, and avoiding stress and tobacco use.
Table 7: Personalized risk assessment and health recommendations generated by the deep learning-based heart disease prediction system.
The system of recommendations allows clinicians and patients to be proactive in managing cardiovascular health, taking predictive information and converting it into practical interventions. Integrating predictive modeling with customized advice, the framework does not only help to determine high-risk people, but also encourages preventive care and lifestyle changes, which eventually helps to detect and treat chronic heart disease at an early stage.

4.3 Discussion

The current paper suggested an effective machine learning and deep learning-based system of early chronic disease (Chronic Kidney Disease (CKD) and chronic heart disease) prediction. The system incorporated state-of-the-art preprocessing, feature engineering, classification models, deep learning systems, risk evaluation, and recommendation systems based on risk assessment to enhance the accuracy of prediction and aid in healthcare decision-making.
To predict CKD, various machine learning models such as Naive Bayes, K-Nearest Neighbors, Decision Tree, Logistic Regression, Support Vector Machine, Random Forest and a Hybrid Ensemble model were run and evaluated. The best overall model was the Random Forest classifier with a perfect value of 1.0000 in terms of accuracy, precision, recall and F1-score. This high performance shows how ensemble learning can efficiently deal with complex nonlinear relationships and heterogeneous clinical data. Random Forest model has been able to reduce errors in classification and at the same time has a high generalization capacity, which is very appropriate in structured medical data. KNN and SVM also showed high predictive power, accuracy and balanced recall rates meaning they can be used to successfully categorize CKD patients.
The CNN-UDRP deep learning model proposed also demonstrated a high predictive accuracy of CKD. Despite the slightly lower accuracy than that of the Random Forest, the CNN-based architecture showed a very strong learning latent semantic and contextual representation of clinical text-like inputs. Hidden dependencies between patient attributes were well represented in the model using embedding and convolutional operations, which makes it very applicable to real-world healthcare settings based on semi-structured or unstructured clinical narratives. The values of the balanced precision, recall, and F1-score show that the model did not change its performance significantly even when there was the imbalance between the classes. Thus, although the numerical performance of Random Forest was the best, CNN-UDRP model offered more scalability and adaptability in intelligent healthcare systems.
To predict chronic heart disease, two deep learning models were used: a standard Deep Artificial Neural Network (ANN) and an Advanced Deep Neural Network (DNN) with an activation of LeakyReLU. Both the models showed very high predictive performance, but the Advanced DNN with LeakyReLU was better than the standard ANN as it had a higher accuracy and less loss values. This enhanced performance can be explained by the fact that the LeakyReLU activation was used and it was effective in countering the dying ReLU problem and also improved the gradient propagation in the training process. Also, batch normalization and dropout regularization enhanced the stability of the model and minimized overfitting. The ideal precision, recall, and F1-score scores of the two models attest to the ability of the models to detect cases of heart disease with high accuracy and minimum false positives. The Advanced DNN was thus found to be the best deep learning system to predict heart disease in this research.
The research was effective in meeting all the set research objectives. The initial goal that aimed at designing and developing disease classification algorithms with machine learning methods was realized by implementing and testing several machine learning and deep learning models in predicting CKD and heart diseases. The latter objective, which was connected with the assessment of disease risk with the help of big data analytics, was achieved due to the implementation of the probabilistic risk stratification mechanisms that divided the patients into low-risk, medium-risk, and high-risk groups according to the probability of prediction. The third goal that was to simplify machine learning methods to be useful in chronic disease prediction was achieved by carrying out rigorous preprocessing, feature selection, dimensionality reduction, and comparative model analysis that yielded highly precise predictive models. Lastly, the fourth goal was met by developing customized recommendation systems that created disease-specific prevention recommendations and clinical suggestions based on predicted risk levels, the proposed framework has a high potential of real-world healthcare implementation through a combination of accurate predictive analytics and personalized clinical recommendations. Combining machine learning, deep learning, and risk-based recommendation systems provides greater early disease detection, helps to implement preventive healthcare approaches, and encourages informed clinical decision-making in chronic disease management.

5. CONCLUSION

This paper proposed an effective machine learning and deep learning-based model to predict chronic diseases (Chronic Kidney Disease (CKD) and chronic heart disease) in their early phases. The framework proposed combined the state-of-the-art preprocessing methods, feature selection, dimensionality reduction, machine learning classifiers, deep learning models, risk evaluation systems, and customized recommendation systems to facilitate intelligent healthcare decision-making. Several machine learning models were tested to predict CKD, and the use of the Random Forest classifier was found to be the most effective predictor because it can effectively address nonlinear clinical and heterogeneous medical data. Similar to the proposed CNN-UDRP model, which demonstrated a high potential to learn latent semantic and contextual representations in clinical narratives, it is applicable in scalable healthcare settings. In predicting chronic heart disease, both deep learning models demonstrated high classification performance with the Advanced Deep Neural Network with LeakyReLU being better than the regular ANN due to better gradient learning and less overfitting. Also, risk stratification and customized recommendation systems were included into the framework, which supplemented the practical utility of the framework by contributing to preventive care and early intervention. In general, the suggested framework has a high potential of clinical applications in real-life practice due to its ability to integrate the accurate disease prediction, intelligent risk assessment, and personalized healthcare recommendations to manage chronic diseases.
In comparasion ,Existing studies mainly focused on single-disease prediction, traditional machine learning optimization, or accuracy enhancement without integrating intelligent clinical support systems. The proposed research will fill these gaps by integrating machine learning and deep learning models to predict CKD and heart disease, PCA-based feature extraction, CNN-UDRP architecture, risk stratification and customized recommendation systems, to deliver a scalable and accurate healthcare decision support model that can be used in clinical practice.

References

  1. A. Ashwini, V. Chirchi, S. Balasubramaniam, and M. A. Shah, “Bio inspired optimization techniques for disease detection in deep learning systems,” Sci. Rep., vol. 15, no. 1, pp. 1–30, 2025, doi: 10.1038/s41598-025-02846-7.
  2. A. I. Torre-Bastida, J. Díaz-de-Arcaya, E. Osaba, K. Muhammad, D. Camacho, and J. Del Ser, Bio-inspired computation for big data fusion, storage, processing, learning and visualization: state of the art and future directions, vol. 37, no. 28. 2025. doi: 10.1007/s00521-021-06332-9.
  3. T. Banerjee and İ. Paçal, “A systematic review of machine learning in heart disease prediction,” Turkish J. Biol., vol. 49, no. SI-5, pp. 600–634, 2025, doi: 10.55730/1300-0152.2766.
  4. R. Islam, A. Sultana, and M. R. Islam, “A comprehensive review for chronic disease prediction using machine learning algorithms,” J. Electr. Syst. Inf. Technol., vol. 11, no. 1, 2024, doi: 10.1186/s43067-024-00150-4.
  5. S. Tutun et al., “An AI-based Decision Support System for Predicting Mental Health Disorders,” Inf. Syst. Front., vol. 25, no. 3, pp. 1261–1276, 2023, doi: 10.1007/s10796-022-10282-5.
  6. A. A. Almazroi, E. A. Aldhahri, S. Bashir, and S. Ashfaq, “A Clinical Decision Support System for Heart Disease Prediction Using Deep Learning,” IEEE Access, vol. 11, no. June, pp. 61646–61659, 2023, doi: 10.1109/ACCESS.2023.3285247.
  7. M. Gholamzadeh, H. Abtahi, and R. Safdari, “The Application of Knowledge-Based Clinical Decision Support Systems to Enhance Adherence to Evidence-Based Medicine in Chronic Disease,” J. Healthc. Eng., vol. 2023, 2023, doi: 10.1155/2023/8550905.
  8. M. Brkan, “Artificial Intelligence and Judicial Decision-Making,” Eur. Data Prot. Law Rev., vol. 9, no. 3, pp. 290–295, 2023, doi: 10.21552/edpl/2023/3/5.
  9. J. Chung and J. Teo, “Mental Health Prediction Using Machine Learning: Taxonomy, Applications, and Challenges,” Appl. Comput. Intell. Soft Comput., vol. 2022, 2022, doi: 10.1155/2022/9970363.
  10. V. Bansla, “A Machine Learning Framework for Early Prediction of Chronic Diseases,” SAMRIDDHI A J. Phys. Sci. Eng. Technol., vol. 17, no. 02, pp. 4–7, 2025, doi: 10.18090/samriddhi.v17i02.02.
  11. M. Harishbhai Tilala et al., “Ethical Considerations in the Use of Artificial Intelligence and Machine Learning in Health Care: A Comprehensive Review,” Cureus, vol. 16, no. 6, 2024, doi: 10.7759/cureus.62443.
  12. S. A. Marrah et al., “Federated deep learning for privacy-preserving disease detection in IoT-enabled healthcare systems,” Front. Comput. Sci., vol. 8, no. February, 2026, doi: 10.3389/fcomp.2026.1725597.
  13. “A Machine Learning – Driven Health Risk Index for Predicting Chronic Disease Burden,” no. April, 2026, doi: 10.14293/PR2199.003245.v2.
  14. D. V. S. Kumar, R. Chaurasia, A. Misra, P. K. Misra, and A. Khang, “Heart Disease and Liver Disease Prediction Using Machine Learning,” Data-Centric AI Solut. Emerg. Technol. Healthc. Ecosyst., vol. 14, no. 04, pp. 205–214, 2023, doi: 10.1201/9781003356189-13.
  15. R. Alanazi, “Identification and Prediction of Chronic Diseases Using Machine Learning Approach,” J. Healthc. Eng., vol. 2022, 2022, doi: 10.1155/2022/2826127.
  16. Z. Xia, P. Xu, Y. Xiong, Y. Lai, and Z. Huang, “Survival Prediction in Patients with Hypertensive Chronic Kidney Disease in Intensive Care Unit: A Retrospective Analysis Based on the MIMIC-III Database,” J. Immunol. Res., vol. 2022, 2022, doi: 10.1155/2022/3377030.
  17. D. S. Khafaga, N. Khodadadi, E. Khodadadi, A. Ali Alhussan, M. M. Eid, and E. S. M. El-Kenawy, “Enhanced early chronic kidney disease prediction using hybrid waterwheel plant algorithm for deep neural network optimization,” Sci. Rep., vol. 15, no. 1, pp. 1–23, 2025, doi: 10.1038/s41598-025-26382-6.
  18. S. K. Ghosh, N. Widatalla, and A. H. Khandoker, “Machine Learning Framework for Early Detection of Chronic Kidney Disease Stages Using Optimized Estimated Glomerular Filtration Rate,” IEEE Access, vol. 13, no. February, pp. 78057–78072, 2025, doi: 10.1109/ACCESS.2025.3565549.
  19. D. Saif, A. M. Sarhan, and N. M. Elshennawy, “Early prediction of chronic kidney disease based on ensemble of deep learning models and optimizers,” J. Electr. Syst. Inf. Technol., vol. 11, no. 1, 2024, doi: 10.1186/s43067-024-00142-4.
  20. S. Dey Tusar et al., “Advancing Chronic Kidney Disease Prediction Through Machine Learning and Deep Learning With Feature Analysis,” Front. Heal. Informatics, vol. 13, no. 3, pp. 11338–11348, 2024, [Online]. Available: https://doi.org/10.52783/fhi.vi.1079
  21. X. Wang et al., “Securing Federated Learning With Blockchain in the Medical Field: Systematic Literature Review,” J. Med. Internet Res., vol. 28, 2026, doi: 10.2196/79052.
  22. S. K. Murarka, N. Jain, V. Sharma, A. Malhosia, and V. K. Chaurasiya, “ML and AI Predictive Modeling and Continuous Patient Surveillance for Chronic Disease Prevention Using Electronic Health Records and IoT-Based Data Streams,” Int. J. Drug Deliv. Technol., vol. 16, no. 8, pp. 138–153, 2026, doi: 10.25258/ijddt.16.8s.17.
  23. L. Zhang et al., “A Novel Prediction Model for Multimodal Medical Data Based on Graph Neural Networks,” Mach. Learn. Knowl. Extr., vol. 7, no. 3, pp. 1–17, 2025, doi: 10.3390/make7030092.
  24. E. Afrifa-Yamoah et al., “Pathways to chronic disease detection and prediction: Mapping the potential of machine learning to the pathophysiological processes while navigating ethical challenges,” Chronic Dis. Transl. Med., vol. 11, no. 1, pp. 1–21, 2025, doi: 10.1002/cdt3.137.
  25. M. J. Sci Ind Res, V. Bhatnagar, S. Sharma, S. Nisha Bhagirath, and D. Sharma, “Mangalayatan Journal of A Machine Learning Framework for Chronic Kidney Disease Analysis Using ORANGE Tool,” vol. 1, no. 1, 2024.
  26. D. Saif, A. M. Sarhan, and N. M. Elshennawy, “Deep-kidney: an effective deep learning framework for chronic kidney disease prediction,” Heal. Inf. Sci. Syst., vol. 12, no. 1, pp. 1–22, 2024, doi: 10.1007/s13755-023-00261-8.
  27. J. Zhu, Y. Wang, S. Wang, and J. Zhou, “Development and validation of a BMI stratified mortality prediction model for patients with COPD complicated by HF using the MIMIC-IV database,” Sci. Rep., vol. 15, no. 1, pp. 1–15, 2025, doi: 10.1038/s41598-025-09605-8.
  28. M. Elhaddad and S. Hamam, “AI-Driven Clinical Decision Support Systems: An Ongoing Pursuit of Potential,” Cureus, vol. 16, no. 4, 2024, doi: 10.7759/cureus.57728.
  29. M. S. Al Reshan et al., “An Innovative Ensemble Deep Learning Clinical Decision Support System for Diabetes Prediction,” IEEE Access, vol. 12, no. March, pp. 106193–106210, 2024, doi: 10.1109/ACCESS.2024.3436641.
  30. J. L. Cross, M. A. Choma, and J. A. Onofrey, “Bias in medical AI: Implications for clinical decision-making,” PLOS Digit. Heal., vol. 3, no. 11, pp. 1–19, 2024, doi: 10.1371/journal.pdig.0000651.
  31. J. N. Saxena, D. V. P. Potturu, and A. Nagraj, “A Hybrid Machine Learning Framework for Early Prediction of Chronic Kidney Disease Progression Using Longitudinal Claims Data : An XGBoost- LSTM Ensemble with Temporal Attention,” 2026.
  32. M. Al Ghamdi and S. Alyahyan, “Enhanced Chronic Kidney Disease Detection: A Hybrid Deep Learning Framework Using Clinical Biomarkers and Ensemble Feature Engineering with DeepCKD-Net,” Appl. Sci., vol. 16, no. 6, pp. 1–32, 2026, doi: 10.3390/app16063024.
  33. S. T. Miller, K. A. Logan, R. Anderson, P. E. Cowell, C. Busby-Earle, and L. D. Morris, “Multi-perspective machine learning MPML: A high-performance and interpretable ensemble method for heart disease prediction,” Mach. Learn. with Appl., vol. 23, p. 100836, 2026, doi: 10.1016/j.mlwa.2026.100836.
  34. R. P. Bhatta and A. Husain, “Analysis of Machine Learning Models for Heart Disease Prediction and Diagnosis NPRC Journal of Multidisciplinary Research,” vol. 3, no. 1, pp. 63–79, 2026, doi: 10.3126/nprcjmr.v3i1.90028.
  35. A. K. Parashar, A. Jamliya, S. Nasrat, R. Soni, H. Disease, and P. Achieving, “XGBoost for Heart Disease Prediction Achieving High Accuracy with Robust Machine Learning Techniques,” pp. 185–191, 2025, doi: 10.69968/ijisem.2025v4i3185-191.
  36. M. Ahmed and I. Husien, “Heart Disease Prediction Using Hybrid Machine Learning: A Brief Review,” J. Robot. Control, vol. 5, no. 3, pp. 884–892, 2024, doi: 10.18196/jrc.v5i3.21606.
  37. G. Manikandan, B. Pragadeesh, V. Manojkumar, A. L. Karthikeyan, R. Manikandan, and A. H. Gandomi, “Classification models combined with Boruta feature selection for heart disease prediction,” Informatics Med. Unlocked, vol. 44, no. January, p. 101442, 2024, doi: 10.1016/j.imu.2023.101442.
  38. N. N. Reddy, L. Nipun, U. Baba, and N. Rishindra, “Optimizing heart disease prediction through ensemble and hybrid machine learning techniques,” vol. 14, no. 5, pp. 5744–5754, 2024, doi: 10.11591/ijece.v14i5.pp5744-5754.
  39. K. Sumwiza, C. Twizere, G. Rushingabigwi, and P. Bakunzibake, “Informatics in Medicine Unlocked Enhanced cardiovascular disease prediction model using random forest algorithm,” Informatics Med. Unlocked, vol. 41, no. March, p. 101316, 2023, doi: 10.1016/j.imu.2023.101316.
  40. L. Kumar, C. Anitha, V. N. Ghodke, N. Nithya, and V. A. Drave, “EAI Endorsed Transactions Deep Learning Based Healthcare Method for Effective Heart Disease Prediction,” vol. 9, 2023, doi: 10.4108/eetpht.9.4283.