BEGIN:VCALENDAR
VERSION:2.0
PRODID:Linklings LLC
BEGIN:VTIMEZONE
TZID:America/New_York
X-LIC-LOCATION:America/New_York
BEGIN:DAYLIGHT
TZOFFSETFROM:-0500
TZOFFSETTO:-0400
TZNAME:EDT
DTSTART:19700308T020000
RRULE:FREQ=YEARLY;BYMONTH=3;BYDAY=2SU
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0400
TZOFFSETTO:-0500
TZNAME:EST
DTSTART:19701101T020000
RRULE:FREQ=YEARLY;BYMONTH=11;BYDAY=1SU
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
DTSTAMP:20210402T160556Z
LOCATION:Track 11
DTSTART;TZID=America/New_York:20201112T160000
DTEND;TZID=America/New_York:20201112T163000
UID:submissions.supercomputing.org_SC20_sess219_ws_prot105@linklings.com
SUMMARY:Exascale Potholes for HPC: Execution Performance and Variability A
 nalysis of the Flagship Application Code HemeLB
DESCRIPTION:Workshop\n\nExascale Potholes for HPC: Execution Performance a
 nd Variability Analysis of the Flagship Application Code HemeLB\n\nWylie\n
 \nPerformance measurement and analysis of parallel applications is often c
 hallenging, despite many excellent commercial and open-source tools being 
 available. Currently envisaged exascale computer systems exacerbate matter
 s by requiring extremely high scalability to effectively exploit millions 
 of processor cores. Unfortunately, significant application execution perfo
 rmance variability arising from increasingly complex interactions between 
 hardware and system software makes this situation much more difficult for 
 application developers and performance analysts alike.  This work consider
 s the performance assessment of the HemeLB exascale-flagship application c
 ode from the EU HPC Centre of Excellence (CoE) CompBioMed running on the S
 uperMUC-NG Tier-0 leadership HPC system, using the methodology of the Perf
 ormance Optimization and Productivity (POP) CoE.  Although 80% scaling eff
 iciency is maintained to over 100,000 MPI processes, disappointing initial
  performance with more processes and corresponding poor strong scaling was
  identified to originate from the same few compute nodes in multiple runs,
  which later system diagnostic checks found had faulty DIMMs and lackluste
 r performance. Excluding these compute nodes from subsequent runs improved
  performance of executions with over 300,000 MPI processes by a factor of 
 five, resulting in 190x speed-up compared to 864 MPI processes.  While com
 munication efficiency remains very good up to the largest scale, parallel 
 efficiency is primarily limited by load balance found to be largely due to
  core-to-core and run-to-run variability from excessive stalls for memory 
 accesses, which affect many HPC systems with Intel Xeon Scalable processor
 s. The POP methodology for this performance diagnosis is demonstrated via 
 a detailed exposition with widely deployed standard measurement and analys
 is tools.\n\nRegistration Category: Workshop Reg Pass
END:VEVENT
END:VCALENDAR

