BEGIN:VCALENDAR
VERSION:2.0
PRODID:Linklings LLC
BEGIN:VTIMEZONE
TZID:America/New_York
X-LIC-LOCATION:America/New_York
BEGIN:DAYLIGHT
TZOFFSETFROM:-0500
TZOFFSETTO:-0400
TZNAME:EDT
DTSTART:19700308T020000
RRULE:FREQ=YEARLY;BYMONTH=3;BYDAY=2SU
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0400
TZOFFSETTO:-0500
TZNAME:EST
DTSTART:19701101T020000
RRULE:FREQ=YEARLY;BYMONTH=11;BYDAY=1SU
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
DTSTAMP:20210402T160554Z
LOCATION:Track 2
DTSTART;TZID=America/New_York:20201112T164500
DTEND;TZID=America/New_York:20201112T171500
UID:submissions.supercomputing.org_SC20_sess208_ws_pmbsf119@linklings.com
SUMMARY:Performance Tradeoffs in GPU Communication: A Study of Host and De
 vice-Initiated Approaches
DESCRIPTION:Workshop\n\nPerformance Tradeoffs in GPU Communication: A Stud
 y of Host and Device-Initiated Approaches\n\nGroves, Brock, Chen, Ibrahim,
  Oliker...\n\nNetwork communication on GPU-based systems is a significant 
 roadblock for many applications with small but frequent messaging requirem
 ents.  One common question for application developers is,  'How can they r
 educe the overheads and achieve the best communication performance on GPUs
 ?'  This work examines device initiated versus host initiated inter-node G
 PU communication using NVSHMEM.  We derive basic communication model param
 eters for single message and batched communication before validating our m
 odel against distributed GEMM benchmarks.  We use our model to estimate pe
 rformance benefits for applications transitioning from CPUs to GPUS for fi
 xed-size and scaled workloads and provide general guidelines for reducing 
 communication overheads.  Our findings show that the host-initiated approa
 ch generally outperforms the device-initiated approach for the system eval
 uated.\n\nRegistration Category: Workshop Reg Pass
END:VEVENT
END:VCALENDAR

