TY - GEN
T1 - Implementing a blocked Aasen's algorithm with a dynamic scheduler on multicore architectures
AU - Ballard, Grey
AU - Becker, Dulceneia
AU - Demmel, James
AU - Dongarra, Jack
AU - Druinsky, Alex
AU - Peled, Inon
AU - Schwartz, Oded
AU - Toledo, Sivan
AU - Yamazaki, Ichitaro
PY - 2013
Y1 - 2013
N2 - Factorization of a dense symmetric indefinite matrix is a key computational kernel in many scientific and engineering simulations. However, there is no scalable factorization algorithm that takes advantage of the symmetry and guarantees numerical stability through pivoting at the same time. This is because such an algorithm exhibits many of the fundamental challenges in parallel programming like irregular data accesses and irregular task dependencies. In this paper, we address these challenges in a tiled implementation of a blocked Aasen's algorithm using a dynamic scheduler. To fully exploit the limited parallelism in this left-looking algorithm, we study several performance enhancing techniques, e.g., parallel reduction to update a panel, tall-skinny LU factorization algorithms to factorize the panel, and a parallel implementation of symmetric pivoting. Our performance results on up to 48 AMD Opteron processors demonstrate that our implementation obtains speedups of up to 2.8 over MKL, while losing only one or two digits in the computed residual norms.
AB - Factorization of a dense symmetric indefinite matrix is a key computational kernel in many scientific and engineering simulations. However, there is no scalable factorization algorithm that takes advantage of the symmetry and guarantees numerical stability through pivoting at the same time. This is because such an algorithm exhibits many of the fundamental challenges in parallel programming like irregular data accesses and irregular task dependencies. In this paper, we address these challenges in a tiled implementation of a blocked Aasen's algorithm using a dynamic scheduler. To fully exploit the limited parallelism in this left-looking algorithm, we study several performance enhancing techniques, e.g., parallel reduction to update a panel, tall-skinny LU factorization algorithms to factorize the panel, and a parallel implementation of symmetric pivoting. Our performance results on up to 48 AMD Opteron processors demonstrate that our implementation obtains speedups of up to 2.8 over MKL, while losing only one or two digits in the computed residual norms.
UR - https://www.scopus.com/pages/publications/84884857876
U2 - 10.1109/IPDPS.2013.98
DO - 10.1109/IPDPS.2013.98
M3 - ???researchoutput.researchoutputtypes.contributiontobookanthology.conference???
AN - SCOPUS:84884857876
T3 - Proceedings - IEEE 27th International Parallel and Distributed Processing Symposium, IPDPS 2013
SP - 895
EP - 907
BT - Proceedings - IEEE 27th International Parallel and Distributed Processing Symposium, IPDPS 2013
PB - IEEE Computer Society
T2 - 27th IEEE International Parallel and Distributed Processing Symposium, IPDPS 2013
Y2 - 20 May 2013 through 24 May 2013
ER -