Workshop:FTXS: Workshop on Fault-Tolerance for HPC at Extreme Scale
Authors: Mohit Kumar and Christian Engelmann (Oak Ridge National Laboratory)
Abstract: Resilience plays an important role in supercomputers by providing correct and efficient operation in case of faults, errors, and failures. Resilience design patterns offer blueprints for effectively applying resilience technologies. Prior work focused on developing initial efficiency and performance models for resilience design patterns. This paper extends it by (1) describing performance, reliability, and availability models for all structural resilience design patterns, (2) providing more detailed models that include flowcharts and state diagrams, and (3) introducing the Resilience Design Pattern Modeling (RDPM) tool that calculates and plots the performance, reliability, and availability metrics of individual patterns and pattern combinations.
Back to FTXS: Workshop on Fault-Tolerance for HPC at Extreme Scale Archive Listing