BEGIN:VCALENDAR
VERSION:2.0
PRODID:Linklings LLC
BEGIN:VTIMEZONE
TZID:America/Chicago
X-LIC-LOCATION:America/Chicago
BEGIN:DAYLIGHT
TZOFFSETFROM:-0600
TZOFFSETTO:-0500
TZNAME:CDT
DTSTART:19700308T020000
RRULE:FREQ=YEARLY;BYMONTH=3;BYDAY=2SU
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0500
TZOFFSETTO:-0600
TZNAME:CST
DTSTART:19701101T020000
RRULE:FREQ=YEARLY;BYMONTH=11;BYDAY=1SU
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
DTSTAMP:20181221T160731Z
LOCATION:C141/143/149
DTSTART;TZID=America/Chicago:20181115T140000
DTEND;TZID=America/Chicago:20181115T143000
UID:submissions.supercomputing.org_SC18_sess192_pap244@linklings.com
SUMMARY:Fault Tolerant One-Sided Matrix Decompositions on Heterogeneous Sy
 stems with GPUs
DESCRIPTION:Paper\nAlgorithms, Architectures, GPUs, Linear Algebra, Networ
 ks, Resiliency, Tech Program Reg Pass\n\nFault Tolerant One-Sided Matrix D
 ecompositions on Heterogeneous Systems with GPUs\n\nChen, Li, Li, Liang, W
 u...\n\nCurrent algorithm-based fault tolerance (ABFT) approach for one-si
 ded matrix decomposition on heterogeneous systems with GPUs have following
  limitations: (1) they do not provide sufficient protection as most of the
 m only maintain checksum in one dimension; (2) their checking scheme is no
 t efficient due to redundant checksum verifications; (3) they fail to prot
 ect PCIe communication; (4) the checksum calculation based on a special ty
 pe of matrix multiplication is far from efficient. By overcoming the above
  limitations, we design an efficient ABFT approach providing stronger prot
 ection for one-sided matrix decomposition methods on heterogeneous systems
 . First, we provide full matrix protection by using checksums in two dimen
 sions. Second, our checking scheme is more efficient by prioritizing the c
 hecksum verification according to the sensitivity of matrix operations to 
 soft errors. Third, we protect PCIe communication by reordering checksum v
 erifications and decomposition steps. Fourth, we accelerate the checksum c
 alculation by 1.7x via better utilizing GPUs.
URL:https://sc18.supercomputing.org/presentation/?id=pap244&sess=sess192
END:VEVENT
END:VCALENDAR

