BEGIN:VCALENDAR
VERSION:2.0
PRODID:Linklings LLC
BEGIN:VTIMEZONE
TZID:America/New_York
X-LIC-LOCATION:America/New_York
BEGIN:DAYLIGHT
TZOFFSETFROM:-0500
TZOFFSETTO:-0400
TZNAME:EDT
DTSTART:19700308T020000
RRULE:FREQ=YEARLY;BYMONTH=3;BYDAY=2SU
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0400
TZOFFSETTO:-0500
TZNAME:EST
DTSTART:19701101T020000
RRULE:FREQ=YEARLY;BYMONTH=11;BYDAY=1SU
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
DTSTAMP:20210402T160558Z
LOCATION:Track 6
DTSTART;TZID=America/New_York:20201113T110500
DTEND;TZID=America/New_York:20201113T113500
UID:submissions.supercomputing.org_SC20_sess225_ws_h2rc105@linklings.com
SUMMARY:Exploring the Acceleration of Nekbone on Reconfigurable Architectu
 res
DESCRIPTION:Workshop\n\nExploring the Acceleration of Nekbone on Reconfigu
 rable Architectures\n\nBrown\n\nHardware technological advances are strugg
 ling to match scientific ambition, and a key question is how we can use th
 e transistors that we already have more effectively. This is especially tr
 ue for HPC, where the tendency is often to throw computation at a problem 
 whereas codes themselves are commonly bound, at-least to some extent, by o
 ther factors. By redesigning an algorithm and moving from a Von Neumann to
  a dataflow style, there is potentially more opportunity to address these 
 bottlenecks on reconfigurable architectures, compared to more general-purp
 ose architectures.\n\nIn this paper we explore the porting of Nekbone’s AX
  kernel, a widely popular HPC mini-app, to FPGAs using high level synthesi
 s via Vitis. While computation is an important part of this code, it is al
 so memory bound on CPUs, and a key question is whether one can ameliorate 
 this by leveraging FPGAs. We first explore optimization strategies for obt
 aining good performance, with over a 4000x runtime difference between the 
 first and final version of our kernel on FPGAs. Subsequently, performance 
 and energy efficiency of our approach on an Alveo U280 are compared agains
 t a 240-core Xeon Platinum CPU and NVIDIA V100 GPU, with the FPGA outperfo
 rming the CPU by around 4x, achieving almost three quarters the GPU perfor
 mance, and displaying significantly more energy efficiency than both. The 
 result of this work is a comparison and a set of techniques that apply to 
 Nekbone on FPGAs specifically and are also of interest more widely in acce
 lerating HPC codes on reconfigurable architectures.\n\nTag: Accelerators, 
 FPGA, and GPUs, Architectures, Emerging Technologies, Heterogeneous System
 s, Reconfigurable Computing\n\nRegistration Category: Workshop Reg Pass
END:VEVENT
END:VCALENDAR

