BEGIN:VCALENDAR
VERSION:2.0
PRODID:Linklings LLC
BEGIN:VTIMEZONE
TZID:America/New_York
X-LIC-LOCATION:America/New_York
BEGIN:DAYLIGHT
TZOFFSETFROM:-0500
TZOFFSETTO:-0400
TZNAME:EDT
DTSTART:19700308T020000
RRULE:FREQ=YEARLY;BYMONTH=3;BYDAY=2SU
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0400
TZOFFSETTO:-0500
TZNAME:EST
DTSTART:19701101T020000
RRULE:FREQ=YEARLY;BYMONTH=11;BYDAY=1SU
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
DTSTAMP:20210402T160100Z
LOCATION:Track 3
DTSTART;TZID=America/New_York:20201118T140000
DTEND;TZID=America/New_York:20201118T143000
UID:submissions.supercomputing.org_SC20_sess158_pap476@linklings.com
SUMMARY:Newton-ADMM: A Distributed GPU-Accelerated Optimizer for Multiclas
 s Classification Problems
DESCRIPTION:Paper\n\nNewton-ADMM: A Distributed GPU-Accelerated Optimizer 
 for Multiclass Classification Problems\n\nFang, Kylasa, Roosta, Mahoney, G
 rama\n\nFirst-order optimization techniques, such as stochastic gradient d
 escent (SGD) and its variants, are widely used in machine learning applica
 tions due to their simplicity and low per-iteration costs. They often requ
 ire, however, large numbers of iterations, with associated communication c
 osts in distributed environments. In contrast, Newton-type methods, while 
 having higher per-iteration computation costs, typically require a signifi
 cantly smaller number of iterations, which directly translates to reduced 
 communication costs.\n\nWe present a novel distributed optimizer for class
 ification problems, which integrates a GPU-accelerated Newton-type solver 
 with the global consensus formulation of Alternating Direction of Method M
 ultipliers (ADMM). By leveraging the communication efficiency of ADMM, a h
 ighly efficient GPU-accelerated inexact-Newton solver, and an effective sp
 ectral penalty parameter selection strategy, we show that our proposed met
 hod: (i) yields better generalization performance on several classificatio
 n problems; (ii) significantly outperforms state-of-the-art methods in dis
 tributed time to solution; and (iii) offers better scaling on large distri
 buted platforms.\n\nTag: Accelerators, FPGA, and GPUs, Algorithms, Graph A
 lgorithms, Scalable Computing\n\nRegistration Category: Tech Program Reg P
 ass
END:VEVENT
END:VCALENDAR

