⚠ Official Notice: www.ijisrt.com is the official website of the International Journal of Innovative Science and Research Technology (IJISRT) Journal for research paper submission and publication. Please beware of fake or duplicate websites using the IJISRT name.



A GAN-Enhanced and Cluster-Aware Data Preprocessing Framework for Robust Predictive Healthcare Analytics


Authors : Aminu Usman Jibril; A. Senthil Kumar

Volume/Issue : Volume 11 - 2026, Issue 7 - July


Google Scholar : https://tinyurl.com/34uzsr7e

Scribd : https://tinyurl.com/ytrezhaf

DOI : https://doi.org/10.38124/ijisrt/26jul1418

Note : A published paper may take 4-5 working days from the publication date to appear in PlumX Metrics, Semantic Scholar, and ResearchGate.


Abstract : Missing values and significant class imbalance are common characteristics of healthcare datasets which significantly impair predictive model performance and reduce their dependability in clinical decision-making. Creation of reliable and broadly applicable healthcare prediction system depends on addressing this issues As to improve overall data quality, this study suggests an integrated data preparation system that integrates cluster-aware oversampling methods with Generative Adversarial Imputation Networks (GAIN). By using adversarial training to understand intricate underlying data distributions GAIN model are used to estimate missing values while maintaining significant statistical correlations between variables. Simultaneously, hybrid SMOTE-ENN method are used to remove ambiguous and noisy data and efficiently handle class imbalance. Real-world diabetic readmission dataset are used to assess suggested methodology, and show notable gains in data completeness distribution preservation, and prediction performance. Significant improvements in accuracy, recall, and F1-score are revealed by experimental data, suggesting improved capacity to detect high-risk individuals. As compared to traditional methods incorporation of sophisticated preprocessing technique enhances model resilience and generalisation. This results highlight significance of integrating class balancing technique and intelligent imputation into single framework. Overall, study emphasises how important sophisticated preprocessing are to enhancing clinical applicability, robustness and dependability of predictive healthcare analytics system.

Keywords : GAN; GAIN; Data Imputation; SMOTE-ENN; Class Imbalance; Predictive Healthcare Analytics; Diabetes Readmission.

References :

  1. Goodfellow, I., Pouget‑Abadie, J., Mirza, M., Xu, B., Warde‑Farley, D., Ozair, S., Courville, A., & Bengio, Y. (2014). Generative adversarial nets. Advances in Neural Information Processing System, 27, 2672–2680. https://papers.nips.cc/paper/5423-generative-adversarial-nets
  2. Yoon, J., Jordon, J., & van der Schaar, M. (2018). GAIN: Missing data imputation using generative adversarial nets. Proceedings of 35th International Conference on Machine Learning 5689–5698. http://proceedings.mlr.press/v80/yoon18a.html
  3. Goodfellow, I., et. al., (2020). Generative adversarial networks. Communications of ACM, 63(11), 139–144. https://doi.org/10.1145/3422622
  4. Chawla, N. V., Bowyer, K. W., Hall, L. O., & Kegelmeyer, W. P. (2002). SMOTE: Synthetic minority over-sampling technique. Journal of Artificial Intelligence Research, 16, 321–357. https://doi.org/10.1613/jair.953
  5. Batista, G. E., Prati, R. C., & Monard, M. C. (2004). study of behavior of several methods for balancing machine learning training data. ACM SIGKDD Explorations 6(1), 20–29. https://doi.org/10.1145/1007730.1007735
  6. He, H., & Garcia, E. A. (2009). Learning from imbalanced data. IEEE Transactions on Knowledge and Data Engineering 21(9), 1263–1284. https://doi.org/10.1109/TKDE.2008.239
  7. King G., & Zeng L. (2001). Logistic regression in rare events data. Political Analysis 9(2), 137–163. https://doi.org/10.1093/oxfordjournals.pan.a004868
  8. Chen, T., & Guestrin, C. (2016). XGBoost: scalable tree boosting system. Proceedings of 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining 785–794. https://doi.org/10.1145/2939672.2939785
  9. Steinhubl, S. R., Muse, E. D., & Topol, E. J. (2015). emerging field of mobile health. Science Translational Medicine, 7(283), 283rv3. https://doi.org/10.1126/scitranslmed.aaa3487
  10. Esteva, A., Robicquet A., Ramsundar, B., Kuleshov, V., DePristo, M., Chou, K., Cui, C., Corrado, G., Thrun, S., & Dean, J. (2019). guide to deep learning in healthcare. Nature Medicine, 25, 24–29. https://doi.org/10.1038/s41591-018-0316-z
  11. Rajkomar, A., Oren, E., Chen, K., Dai, A. M., Hajaj, N., Hardt M., Liu, P. J., Liu, X., Marcus J., Sun, M., Sundberg P., Yee, H., Zhang K., Zhang Y., Flores G., Ledsam, J. R., et. al., (2018). Scalable and accurate deep learning with electronic health records. npj Digital Medicine, 1(18). https://doi.org/10.1038/s41746-018-0029-1
  12. Bates J., Saria, S., Ohno-Machado, L., Shah, A., & Escobar, G. (2014). Big data in health care. Health Affairs 33(7), 1123–1131. https://doi.org/10.1377/hlthaff.2014.0147
  13. Creswell, A., White, T., Dumoulin, V., Arulkumaran, K., Sengupta, B., & Bharath, A. A. (2018). Generative adversarial networks: An overview. IEEE Signal Processing Magazine, 35(1), 53–65. https://doi.org/10.1109/MSP.2017.2765202
  14. Breiman, L. (2001). Random forests. Machine Learning 45, 5–32. https://doi.org/10.1023/A:1010933404324
  15. Friedman, J. H. (2001). Greedy function approximation: gradient boosting machine. Annals of Statistics 29(5), 1189–1232. https://doi.org/10.1214/aos/1013203451
  16. LeCun, Y., Bengio, Y., & Hinton, G. (2015). Deep learning. Nature, 521, 436–444. https://doi.org/10.1038/nature14539
  17. Mikolov, T., Chen, K., Corrado, G., & Dean, J. (2013). Efficient estimation of word representations in vector space. ICLR Workshop. https://arxiv.org/abs/1301.3781
  18. Kingma, D. P., & Ba, J. (2015). Adam: method for stochastic optimization. ICLR. https://arxiv.org/abs/1412.6980
  19. Van Buuren, S., & Groothuis-Oudshoorn, K. (2011). MICE: Multivariate imputation by chained equations. Journal of Statistical Software, 45(3), 1–67. https://www.jstatsoft.org/article/view/v045i03
  20. Little, R. J., & Rubin, D. B. (2019). Statistical analysis with missing data (3rd ed.). Wiley. https://www.wiley.com/en-us/Statistical+Analysis+with+Missing+Data%2C+3rd+Edition-p-9781119407605
  21. Dong G., & Peng H. (2013). Principled missing data methods for researchers. Springer. https://doi.org/10.1007/978-1-4614-6841-3
  22. Schafer, J. L. (1997). Analysis of incomplete multivariate data. Chapman & Hall. https://doi.org/10.1007/978-1-4757-3542-0
  23. Fawcett T. (2006). An introduction to ROC analysis. Pattern Recognition Letters 27(8), 861–874. https://doi.org/10.1016/j.patrec.2005.10.010
  24. Chicco, D., & Jurman, G. (2020). advantages of Matthews correlation coefficient (MCC) over F1 score and accuracy in binary classification evaluation. Bioinformatics 36(20), 517–520. https://doi.org/10.1093/bioinformatics/btaa106
  25. Sculley, D., Holt G., Golovin, D., Davydov, E., Phillips T., Ebner, D., Chaudhary, V., Young M., Crespo, J. F., & Dennison, D. (2015). Hidden technical debt in machine learning system. Advances in Neural Information Processing System. https://papers.nips.cc/paper/5656-hidden-technical-debt-in-machine-learning-system
  26. Strack, B., DeShazo, J., Gennings C., Olmo, J. L., Ventura, S., Cios K., & Re, C. (2014). Impact of HbA1c measurement on hospital readmission rates. BioMed Research International, 2014, Article ID 781670. https://doi.org/10.1155/2014/781670
  27. He, H., Bai, Y., Garcia, E., & Li, S. (2008). ADASYN: Adaptive synthetic sampling approach for imbalanced learning. IEEE International Joint Conference on Neural Networks 1322–1328. https://doi.org/10.1109/IJCNN.2008.4633969
  28. Krawczyk, J. (2016). Learning from imbalanced data: open challenges and future directions. Progress in Artificial Intelligence, 5, 221–232. https://doi.org/10.1007/s13748-016-0094-0
  29. Iqbal, Z., Rafique, A., Qaisar, S., et. al., (2025). Advancements and challenges in development of generative adversarial network (GANs) for deep learning. SN Computer Science. https://doi.org/10.1007/s44354-025-00007-w
  30. Ou, H., Yao, Y., & He, Y. (2024). Missing Data Imputation Method Combining Random Forest and Generative Adversarial Imputation Network. Sensors 24(4), 1112. https://doi.org/10.3390/s24041112

Missing values and significant class imbalance are common characteristics of healthcare datasets which significantly impair predictive model performance and reduce their dependability in clinical decision-making. Creation of reliable and broadly applicable healthcare prediction system depends on addressing this issues As to improve overall data quality, this study suggests an integrated data preparation system that integrates cluster-aware oversampling methods with Generative Adversarial Imputation Networks (GAIN). By using adversarial training to understand intricate underlying data distributions GAIN model are used to estimate missing values while maintaining significant statistical correlations between variables. Simultaneously, hybrid SMOTE-ENN method are used to remove ambiguous and noisy data and efficiently handle class imbalance. Real-world diabetic readmission dataset are used to assess suggested methodology, and show notable gains in data completeness distribution preservation, and prediction performance. Significant improvements in accuracy, recall, and F1-score are revealed by experimental data, suggesting improved capacity to detect high-risk individuals. As compared to traditional methods incorporation of sophisticated preprocessing technique enhances model resilience and generalisation. This results highlight significance of integrating class balancing technique and intelligent imputation into single framework. Overall, study emphasises how important sophisticated preprocessing are to enhancing clinical applicability, robustness and dependability of predictive healthcare analytics system.

Keywords : GAN; GAIN; Data Imputation; SMOTE-ENN; Class Imbalance; Predictive Healthcare Analytics; Diabetes Readmission.

Paper Submission Last Date
31 - August - 2026

SUBMIT YOUR PAPER CALL FOR PAPERS
Video Explanation for Published paper

Never miss an update from Papermashup

Get notified about the latest tutorials and downloads.

Subscribe by Email

Get alerts directly into your inbox after each post and stay updated.
Subscribe
OR

Subscribe by RSS

Add our RSS to your feedreader to get regular updates from us.
Subscribe