BEGIN:VCALENDAR
VERSION:2.0
PRODID:Linklings LLC
BEGIN:VTIMEZONE
TZID:America/New_York
X-LIC-LOCATION:America/New_York
BEGIN:DAYLIGHT
TZOFFSETFROM:-0500
TZOFFSETTO:-0400
TZNAME:EDT
DTSTART:19700308T020000
RRULE:FREQ=YEARLY;BYMONTH=3;BYDAY=2SU
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0400
TZOFFSETTO:-0500
TZNAME:EST
DTSTART:19701101T020000
RRULE:FREQ=YEARLY;BYMONTH=11;BYDAY=1SU
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
DTSTAMP:20210402T160030Z
LOCATION:Track 8
DTSTART;TZID=America/New_York:20201118T153400
DTEND;TZID=America/New_York:20201118T155100
UID:submissions.supercomputing.org_SC20_sess355_spostg113@linklings.com
SUMMARY:Memory-Centric 3D Image Reconstruction with Hierarchical Communica
 tions on Multi-GPU Node Architecture
DESCRIPTION:ACM Student Research Competition: Graduate Poster, ACM Student
  Research Competition: Undergraduate Poster\n\nMemory-Centric 3D Image Rec
 onstruction with Hierarchical Communications on Multi-GPU Node Architectur
 e\n\nHidayetoglu\n\nX-ray computed tomography is a commonly used technique
  for non-invasive imaging at synchrotron facilities. Iterative tomographic
  reconstruction algorithms are often preferred for recovering high quality
  3D volumetric images from 2D X-ray images, however, their use has been li
 mited to small/medium datasets due to their computational requirements. In
  this work, we propose a high-performance iterative reconstruction system 
 for terabyte(s)-scale 3D volumes. Our design involves three novel optimiza
 tions: (1) optimization of (back)projection operators by extending the 2D 
 memory-centric approach to 3D; (2) performing hierarchical communications 
 by exploiting “fat-node" architecture with many GPUs; (3) utilization of m
 ixed-precision types while preserving convergence rate and quality. We ext
 ensively evaluate the proposed optimizations and scaling on the Summit sup
 ercomputer. Our largest reconstruction is a mouse brain volume with 9K×11K
 ×11K voxels, where the total reconstruction time is under three minutes us
 ing 24,576 GPUs, reaching 65 PFLOPS: 34% of Summit’s peak performance.\n\n
 Tag: Student Program\n\nRegistration Category: Tech Program Reg Pass
END:VEVENT
END:VCALENDAR

