A Scalable Distributed and Fault-Tolerant Architecture for Cloud-Based Machine Learning and Data Analysis

  • GBOR, Grace Dooshima Department of Computer Science, College of Physical Sciences, Joseph Sarwuan Tarka University, Benue State, Makurdi, Nigeria.
  • Emmanuel Ogala Department of Computer Science, College of Physical Sciences, Joseph Sarwuan Tarka University, Benue State, Makurdi, Nigeria.
  • Donald Douglas Atsa'am Department of Computer Science, College of Physical Sciences, Joseph Sarwuan Tarka University, Benue State, Makurdi, Nigeria.
  • Iorshashe Agaji Department of Computer Science, College of Physical Sciences, Joseph Sarwuan Tarka University, Benue State, Makurdi, Nigeria.
Keywords: Cloud-Based Machine Learning, Scalable Architecture, Fault tolerance, Parallel processing, Resource management

Abstract

The rapid growth of data-intensive applications has necessitated the development of scalable and efficient architectures for cloud-based machine learning and data analysis. This study proposes a scalable, distributed, and fault-tolerant architecture designed to address the challenges of processing large-scale and dynamic datasets in cloud environments. The architecture integrates key components, including data ingestion, distributed storage, parallel processing frameworks, machine learning pipelines, and application deployment layers, enabling seamless data flow and modular system design. It supports both batch and real-time data processing, making it adaptable to diverse analytical workloads. A design science and experimental research methodology was adopted to develop and evaluate the proposed system. Mathematical modeling and performance analysis were employed to assess system scalability, throughput, and latency under varying load conditions. Experimental results demonstrated that the architecture achieves significant improvements in processing efficiency and resource utilization through horizontal scaling. However, the findings also revealed sub-linear scalability behavior due to factors such as communication overhead, synchronization delays, and resource contention, which are inherent in distributed systems. The architecture exhibited strong fault tolerance and resilience, ensuring continuous system operation through redundancy and dynamic resource management. Performance evaluations highlighted an optimal operating region where throughput is maximized and latency remains within acceptable limits, beyond which system performance begins to degrade. The proposed architecture provides a robust and flexible framework for large-scale machine learning and data analysis in cloud environments. It offers a balance between scalability, performance, and reliability, making it suitable for modern data-driven applications. Future research may focus on enhancing auto-scaling strategies, optimizing workload distribution, and incorporating intelligent resource management techniques to further improve system efficiency.

Downloads

Download data is not yet available.

References

Alzoubi, Y. I., Mishra, A., & Topcu, A. E. (2024). Research trends in deep learning and machine learning for cloud computing security. Artificial Intelligence Review, 57, 132.

https://doi.org/10.1007/s10462-024-10776-5

Ambika, B., Pandian, S. M. V., Sundaravadivel, P., & Vinothkumar, E. S. (2026). Federated deep reinforcement learning enabled hierarchical edge–fog–cloud architecture for intelligent task offloading in 6G networks. Scientific Reports.

https://doi.org/10.1038/s41598-026-57660-6

Asghar, A., Rehman, A. U., Ayaz, R., & Suryana, A. (2025). Smart cloud architectures: The combination of machine learning and cloud computing. Engineering Proceedings, 107(1), 74.

https://doi.org/10.3390/engproc2025107074

Chockalingam, N., Deshpande, A., Butra, L., Bodala, R. S., Saksena, N., Parthasarathy, A., & Agarwal, A. K. (2025). Scalable cloud-native architectures for intelligent data processing. arXiv.

https://arxiv.org/abs/2512.22231

Laxkar, P., & Jain, N. (2025). A review of scalable machine learning architectures in cloud environments: Challenges and innovations. International Journal of Scientific Research in Computer Science, Engineering and Information Technology, 11(2), 2907–2916.

https://doi.org/10.32628/CSEIT25112764

Manhary, F. N., Mohamed, M. H., & Farouk, M. (2025). A scalable machine learning strategy for resource allocation in database systems. Scientific Reports, 15, 30567.

https://doi.org/10.1038/s41598-025-14962-5

Mills, N., Issadeen, Z., Matharaarachchi, A., Bandaragoda, T., & De Silva, D. (2024). A cloud-based architecture for explainable big data analytics using self-structuring artificial intelligence. Discover Artificial Intelligence, 4(33).

https://doi.org/10.1007/s44163-024-00123-6

Mungoli, N. (2023). Scalable, distributed AI frameworks: Leveraging cloud computing for enhanced deep learning performance and efficiency. arXiv.

https://arxiv.org/abs/2304.13738

Nashaat, M., Moussa, W., Rizk, R., & Saber, W. (2026). Dynamic machine learning approach for workload prediction in cloud environments. Scientific Reports, 16, 10983.

https://doi.org/10.1038/s41598-026-40777-z

Punniyamoorthy, V., Parthi, A. G., Palanigounder, M., Kodali, R. K., Kumar, B., & Kannan, K. (2025). A privacy-preserving cloud architecture for distributed machine learning at scale. arXiv.

https://arxiv.org/abs/2512.10341

Sanyadanam, A., & Srirama, S. N. (2026). Serverless data pipeline architecture supporting distributed machine learning on fog devices. Computer Communications.

https://doi.org/10.1016/j.comcom.2026.108505

Simaiya, S., Lilhore, U. K., Sharma, Y. K., Rao, K. B. V., & Rao, V. V. R. M. (2024). A hybrid cloud load balancing and host utilization prediction method using deep learning. Scientific Reports, 14, 1337.

https://doi.org/10.1038/s41598-024-51466-0

Wang, Z., Liu, Y., & Huang, J. (2024). An open API architecture to discover trustworthy explanation of cloud AI services. IEEE Transactions on Cloud Computing.

https://doi.org/10.1109/TCC.2024.3398609

Published
2026-07-23
How to Cite
Grace Dooshima, G., Ogala, E., Douglas Atsa’am, D., & Agaji, I. (2026). A Scalable Distributed and Fault-Tolerant Architecture for Cloud-Based Machine Learning and Data Analysis. GPH-International Journal of Computer Science and Engineering, 9(2), 01-21. https://doi.org/10.5281/zenodo.21509041