High-Availability HPC Cluster Design for Mission-Critical Public Infrastructure: Lessons from Energy Grid and Government

Authors

  • Rakesh Challa Principal Engineer, Dell Technologies Author

DOI:

https://doi.org/10.15662/5k48pc50

Keywords:

High-Availability HPC, Fault-Tolerant Clusters, Mission-Critical Infrastructure, Multi-Layer Failover, Redundant Pathing

Abstract

The paper addresses the design issue of fault-tolerant high-availability HPC clusters of mission- critical systems in the public-sector, such as energy grid operators and government data centers. The paper focuses on such architectural designs as redundant pathing, multi-layer failure, homogeneity in the firmware and risk managed cutovers to assist in achieving zero-downtime operations. The system performance was quantitatively evaluated by running a simulation-based experiment which incorporated Monte Carlo fault injection. Multi-layer failover with firmware homogeneity availability, MTTR, and recovery success was 99.97, 38 and 99.9 seconds respectively, whereas controlled cutovers also increased availability to 99.98. The workload was completed over 99.5 percent even with several faults. These findings point to the fact that properly designed HPC systems may be maintained, and used in handling large volumes of data and be cyber- resilient in serving millions of people. The paper provides a methodical method of designing and assessing the high availability HPC systems that can be applied in the national interest activities

References

[1] Egwutuoha, I. P., Levy, D., Selic, B., & Chen, S. (2013). A survey of fault tolerance mechanisms and checkpoint/restart implementations for high performance computing systems. The Journal of Supercomputing, 65(3), 1302–1326. https://doi.org/10.1007/s11227-013-0884-0

[2] Boukerche, A., Al-Shaikh, R. A., & Notare, M. S. M. A. (2007). Towards highly available and scalable high performance clusters. Journal of Computer and System Sciences, 73(8), 1240–1251. https://doi.org/10.1016/j.jcss.2007.02.011

[3] Somasekaram, P., Calinescu, R., & Buyya, R. (2021). High-availability clusters: A taxonomy, survey, and future directions. Journal of Systems and Software, 187, 111208. https://doi.org/10.1016/j.jss.2021.111208

[4] Leangsuksun, C. B., Shen, L., Liu, T., & Scott, S. L. (2004). Achieving high availability and performance computing with an HA-OSCAR cluster. Future Generation Computer Systems, 21(4), 597–606. https://doi.org/10.1016/j.future.2003.12.026

[5] Mesbahi, M. R., Rahmani, A. M., & Hosseinzadeh, M. (2018). Reliability and high availability in cloud computing environments: a reference roadmap. Human-centric Computing and Information Sciences, 8(1). https://doi.org/10.1186/s13673-018-0143-8

[6] Netti, A., Kiziltan, Z., Babaoglu, O., Sirbu, A., Bartolini, A., & Borghesi, A. (2018, October 26). Online Fault Classification in HPC Systems through Machine Learning. arXiv.org. https://arxiv.org/abs/1810.11208

[7] Hukerikar, S., & Engelmann, C. (2017). Resilience Design Patterns: A structured approach to resilience at extreme scale. Supercomputing Frontiers and Innovations, 4(3). https://doi.org/10.14529/jsfi170301

[8] Ashraf, R. A., Hukerikar, S., & Engelmann, C. (2018). Pattern-based Modeling of Multiresilience Solutions for High- Performance Computing. Pattern-based Modeling of Multiresilience Solutions for High-Performance Computing, 80–87. https://doi.org/10.1145/3184407.3184421

[9] Păun, A., Chandler, C., Leangsuksun, C. B., & Păun, M. (2016). A failure index for HPC applications. Journal of Parallel and Distributed Computing, 93–94, 146–153. https://doi.org/10.1016/j.jpdc.2016.04.009

[10] Benacchio, T., Bonaventura, L., Altenbernd, M., Cantwell, C. D., Düben, P. D., Gillard, M., Giraud, L., Göddeke, D., Raffin, E., Teranishi, K., & Wedi, N. (2021). Resilience and fault tolerance in high-performance computing for numerical weather and climate prediction. The International Journal of High Performance Computing Applications, 35(4), 285–311. https://doi.org/10.1177/1094342021990433

[11] Georgakoudis, G., Guo, L., & Laguna, I. (2021). REINIT++: Evaluating the performance of Global-Restart Recovery Methods for MPI fault Tolerance. arXiv (Cornell University), 536–554. https://doi.org/10.48550/arxiv.2102.06896

[12] Bouteiller, A., & Bosilca, G. (2022). Implicit Actions and Non-blocking Failure Recovery with MPI. Implicit Actions and Non-blocking Failure Recovery With MPI, 17, 36–46. https://doi.org/10.1109/ftxs56515.2022.00009

[13] Bharany, S., Badotra, S., Sharma, S., Rani, S., Alazab, M., Jhaveri, R. H., & Gadekallu, T. R. (2022). Energy efficient fault tolerance techniques in green cloud computing: A systematic survey and taxonomy. Sustainable Energy Technologies and Assessments, 53, 102613. https://doi.org/10.1016/j.seta.2022.102613

[14] Moran, M., Balladini, J., Rexachs, D., & Rucci, E. (2020). Towards Management of Energy Consumption in HPC Systems with Fault Tolerance. 2020 IEEE Congreso Bienal De Argentina (ARGENCON), 1– 8. https://doi.org/10.1109/argencon49523.2020.9505498

[15] Iaeme. (2023). DEVELOPMENT OF a HIGH-AVAILABILITY CLUSTER FOR FAULT-TOLERANT COMPUTING. Arabixiv (OSF Preprints). https://doi.org/10.17605/osf.io/2u9qg

[16] Li, Z., Chang, V., Hu, H., Hu, H., Li, C., & Ge, J. (2021). Real-time and dynamic fault-tolerant scheduling for scientific workflows in clouds. Information Sciences, 568, 13–39. https://doi.org/10.1016/j.ins.2021.03.003

[17] Martinez, H. F., Mondragon, O. H., Rubio, H. A., & Marquez, J. (2022). Computational and communication infrastructure challenges for resilient cloud services. Computers, 11(8),118. https://doi.org/10.3390/computers11080118

[18] Castro-León, M., Meyer, H., Rexachs, D., & Luque, E. (2015). Fault tolerance at system level based on RADIC architecture. Journal of Parallel and Distributed Computing, 86, 98–111. https://doi.org/10.1016/j.jpdc.2015.08.005

Downloads

Published

2024-11-09

How to Cite

High-Availability HPC Cluster Design for Mission-Critical Public Infrastructure: Lessons from Energy Grid and Government . (2024). International Journal of Engineering & Extended Technologies Research (IJEETR), 6(6), 9305-9309. https://doi.org/10.15662/5k48pc50