BEGIN:VCALENDAR
VERSION:2.0
PRODID:Linklings LLC
BEGIN:VTIMEZONE
TZID:America/New_York
X-LIC-LOCATION:America/New_York
BEGIN:DAYLIGHT
TZOFFSETFROM:-0500
TZOFFSETTO:-0400
TZNAME:EDT
DTSTART:19700308T020000
RRULE:FREQ=YEARLY;BYMONTH=3;BYDAY=2SU
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0400
TZOFFSETTO:-0500
TZNAME:EST
DTSTART:19701101T020000
RRULE:FREQ=YEARLY;BYMONTH=11;BYDAY=1SU
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
DTSTAMP:20210402T160546Z
LOCATION:Track 9
DTSTART;TZID=America/New_York:20201113T162500
DTEND;TZID=America/New_York:20201113T165500
UID:submissions.supercomputing.org_SC20_sess231_ws_mlcs105@linklings.com
SUMMARY:Explainable Machine Learning Frameworks for Managing HPC Systems
DESCRIPTION:Workshop\n\nExplainable Machine Learning Frameworks for Managi
 ng HPC Systems\n\nAksar, Ateş, Leung, Coskun\n\nRecent research on su
 percomputing proposes a variety of machine learning frameworks that are ab
 le to detect performance variations, find optimum application configuratio
 ns, perform intelligent scheduling or node allocation and improve system s
 ecurity. Although these goals align well with HPC systems' needs, barriers
  such as the lack of user trust or the difficulty of debugging need to be 
 overcome to enable the widespread adoption of such frameworks in productio
 n systems. This paper evaluates a new counterfactual time series explainab
 ility method and compares it against state-of-the-art explainability metho
 ds for supervised machine learning frameworks that use multivariate HPC sy
 stem telemetry data. The counterfactual time series explainability method 
 outperforms existing methods in terms of comprehensibility and robustness.
  We also show how explainability techniques can be used to debug machine l
 earning frameworks and gain a better understanding of HPC system telemetry
  data.\n\nRegistration Category: Workshop Reg Pass
END:VEVENT
END:VCALENDAR

