Matches in SemOpenAlex for { <https://semopenalex.org/work/W3006153736> ?p ?o ?g. }
- W3006153736 endingPage "234" @default.
- W3006153736 startingPage "225" @default.
- W3006153736 abstract "Dense matrix factorizations, such as LU, Cholesky and QR, are widely used for scientific applications that require solving systems of linear equations, eigenvalues and linear least squares problems. Such computations are normally carried out on supercomputers, whose ever-growing scale induces a fast decline of the Mean Time To Failure (MTTF). This paper proposes a new hybrid approach, based on Algorithm-Based Fault Tolerance (ABFT), to help matrix factorizations algorithms survive fail-stop failures. We consider extreme conditions, such as the absence of any reliable component and the possibility of loosing both data and checksum from a single failure. We will present a generic solution for protecting the right factor, where the updates are applied, of all above mentioned factorizations. For the left factor, where the panel has been applied, we propose a scalable checkpointing algorithm. This algorithm features high degree of checkpointing parallelism and cooperatively utilizes the checksum storage leftover from the right factor protection. The fault-tolerant algorithms derived from this hybrid solution is applicable to a wide range of dense matrix factorizations, with minor modifications. Theoretical analysis shows that the fault tolerance overhead sharply decreases with the scaling in the number of computing units and the problem size. Experimental results of LU and QR factorization on the Kraken (Cray XT5) supercomputer validate the theoretical evaluation and confirm negligible overhead, with- and without-errors." @default.
- W3006153736 created "2020-02-24" @default.
- W3006153736 creator A5026991681 @default.
- W3006153736 creator A5029586169 @default.
- W3006153736 creator A5054873210 @default.
- W3006153736 creator A5071431152 @default.
- W3006153736 creator A5076541920 @default.
- W3006153736 date "2012-02-25" @default.
- W3006153736 modified "2023-10-18" @default.
- W3006153736 title "Algorithm-based fault tolerance for dense matrix factorizations" @default.
- W3006153736 cites W2001495258 @default.
- W3006153736 cites W2072072075 @default.
- W3006153736 cites W2083606889 @default.
- W3006153736 cites W2083613288 @default.
- W3006153736 cites W2096504919 @default.
- W3006153736 cites W2151984682 @default.
- W3006153736 cites W2158344138 @default.
- W3006153736 cites W2165009364 @default.
- W3006153736 cites W2296772319 @default.
- W3006153736 cites W4231150350 @default.
- W3006153736 cites W4239025233 @default.
- W3006153736 doi "https://doi.org/10.1145/2370036.2145845" @default.
- W3006153736 hasPublicationYear "2012" @default.
- W3006153736 type Work @default.
- W3006153736 sameAs 3006153736 @default.
- W3006153736 citedByCount "28" @default.
- W3006153736 countsByYear W30061537362013 @default.
- W3006153736 countsByYear W30061537362014 @default.
- W3006153736 countsByYear W30061537362015 @default.
- W3006153736 countsByYear W30061537362016 @default.
- W3006153736 countsByYear W30061537362017 @default.
- W3006153736 countsByYear W30061537362018 @default.
- W3006153736 countsByYear W30061537362019 @default.
- W3006153736 countsByYear W30061537362020 @default.
- W3006153736 countsByYear W30061537362021 @default.
- W3006153736 countsByYear W30061537362023 @default.
- W3006153736 crossrefType "journal-article" @default.
- W3006153736 hasAuthorship W3006153736A5026991681 @default.
- W3006153736 hasAuthorship W3006153736A5029586169 @default.
- W3006153736 hasAuthorship W3006153736A5054873210 @default.
- W3006153736 hasAuthorship W3006153736A5071431152 @default.
- W3006153736 hasAuthorship W3006153736A5076541920 @default.
- W3006153736 hasConcept C106487976 @default.
- W3006153736 hasConcept C111919701 @default.
- W3006153736 hasConcept C11413529 @default.
- W3006153736 hasConcept C120314980 @default.
- W3006153736 hasConcept C121332964 @default.
- W3006153736 hasConcept C123213974 @default.
- W3006153736 hasConcept C158693339 @default.
- W3006153736 hasConcept C159985019 @default.
- W3006153736 hasConcept C162372511 @default.
- W3006153736 hasConcept C173608175 @default.
- W3006153736 hasConcept C188060507 @default.
- W3006153736 hasConcept C192562407 @default.
- W3006153736 hasConcept C2779960059 @default.
- W3006153736 hasConcept C34727166 @default.
- W3006153736 hasConcept C41008148 @default.
- W3006153736 hasConcept C42355184 @default.
- W3006153736 hasConcept C44363057 @default.
- W3006153736 hasConcept C46085209 @default.
- W3006153736 hasConcept C48044578 @default.
- W3006153736 hasConcept C62520636 @default.
- W3006153736 hasConcept C63540848 @default.
- W3006153736 hasConcept C77088390 @default.
- W3006153736 hasConcept C83283714 @default.
- W3006153736 hasConceptScore W3006153736C106487976 @default.
- W3006153736 hasConceptScore W3006153736C111919701 @default.
- W3006153736 hasConceptScore W3006153736C11413529 @default.
- W3006153736 hasConceptScore W3006153736C120314980 @default.
- W3006153736 hasConceptScore W3006153736C121332964 @default.
- W3006153736 hasConceptScore W3006153736C123213974 @default.
- W3006153736 hasConceptScore W3006153736C158693339 @default.
- W3006153736 hasConceptScore W3006153736C159985019 @default.
- W3006153736 hasConceptScore W3006153736C162372511 @default.
- W3006153736 hasConceptScore W3006153736C173608175 @default.
- W3006153736 hasConceptScore W3006153736C188060507 @default.
- W3006153736 hasConceptScore W3006153736C192562407 @default.
- W3006153736 hasConceptScore W3006153736C2779960059 @default.
- W3006153736 hasConceptScore W3006153736C34727166 @default.
- W3006153736 hasConceptScore W3006153736C41008148 @default.
- W3006153736 hasConceptScore W3006153736C42355184 @default.
- W3006153736 hasConceptScore W3006153736C44363057 @default.
- W3006153736 hasConceptScore W3006153736C46085209 @default.
- W3006153736 hasConceptScore W3006153736C48044578 @default.
- W3006153736 hasConceptScore W3006153736C62520636 @default.
- W3006153736 hasConceptScore W3006153736C63540848 @default.
- W3006153736 hasConceptScore W3006153736C77088390 @default.
- W3006153736 hasConceptScore W3006153736C83283714 @default.
- W3006153736 hasIssue "8" @default.
- W3006153736 hasLocation W30061537361 @default.
- W3006153736 hasOpenAccess W3006153736 @default.
- W3006153736 hasPrimaryLocation W30061537361 @default.
- W3006153736 hasRelatedWork W2052455844 @default.
- W3006153736 hasRelatedWork W2059724662 @default.
- W3006153736 hasRelatedWork W2112426221 @default.
- W3006153736 hasRelatedWork W2113085198 @default.
- W3006153736 hasRelatedWork W2123500455 @default.
- W3006153736 hasRelatedWork W2127054029 @default.