BEGIN:VCALENDAR
VERSION:2.0
PRODID:Linklings LLC
BEGIN:VTIMEZONE
TZID:America/New_York
X-LIC-LOCATION:America/New_York
BEGIN:DAYLIGHT
TZOFFSETFROM:-0500
TZOFFSETTO:-0400
TZNAME:EDT
DTSTART:19700308T020000
RRULE:FREQ=YEARLY;BYMONTH=3;BYDAY=2SU
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0400
TZOFFSETTO:-0500
TZNAME:EST
DTSTART:19701101T020000
RRULE:FREQ=YEARLY;BYMONTH=11;BYDAY=1SU
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
DTSTAMP:20210402T160556Z
LOCATION:Track 1
DTSTART;TZID=America/New_York:20201111T111500
DTEND;TZID=America/New_York:20201111T114000
UID:submissions.supercomputing.org_SC20_sess193_ws_hipar103@linklings.com
SUMMARY:A Case Study and Characterization of a Many-Socket, Multi-Tier NUM
 A HPC Platform
DESCRIPTION:Workshop\n\nA Case Study and Characterization of a Many-Socket
 , Multi-Tier NUMA HPC Platform\n\nImes, Hofmeyr, Kang, Walters\n\nAs the n
 umber of processor cores and sockets on HPC compute nodes increase and sys
 tems expose more hierarchical non-uniform memory access (NUMA) architectur
 es, efficiently scaling applications within even a single shared memory sy
 stem is becoming more challenging. It is now common for HPC compute nodes 
 to have two or more sockets and dozens of cores, but future generation sys
 tems may contain an order of magnitude more of each. We conduct experiment
 s on a state-of-the-art Intel Xeon Platinum system with 12 processor socke
 ts, totaling 288 cores (576 hardware threads), arranged in a multi-tier NU
 MA hierarchy. Platforms of this scale and memory hierarchy are uncommon to
 day, providing us a unique opportunity to empirically evaluate—rather than
  model or simulate—an architecture potentially representative of future HP
 C compute nodes. We quantify the platform’s multi-tier NUMA patterns, then
  evaluate its suitability for HPC workloads using a modern HPC metagenome 
 assembler application as a case study, and other HPC benchmarks with a var
 iety of parallelization techniques to characterize the system’s performanc
 e, scalability, I/O patterns, and performance/power behavior. Our results 
 demonstrate near- perfect scaling for embarrassingly parallel and weak sca
 ling workloads, but challenges for random memory access workloads. For the
  latter, we find poor scaling performance with the default scheduling appr
 oaches—e.g., which do not pin threads— suggesting that userspace or kernel
  schedulers may require changes to better manage the multi-tier NUMA hiera
 rchies of very large shared memory platforms.\n\nTag: Extreme Scale Comput
 ing, Heterogeneous Systems, Parallel Programming Languages, Libraries, and
  Models, Portability, Resource Management and Scheduling, Scalable Computi
 ng\n\nRegistration Category: Workshop Reg Pass
END:VEVENT
END:VCALENDAR

