TY - GEN
T1 - Short Paper
T2 - 2023 International Conference on High Performance Computing, Network, Storage, and Analysis, SC Workshops 2023
AU - Aach, Marcel
AU - Sarma, Rakesh
AU - Inanc, Eray
AU - Riedel, Morris
AU - Lintermann, Andreas
N1 - Publisher Copyright: © 2023 ACM.
PY - 2023/11/12
Y1 - 2023/11/12
N2 - Hyperparameter Optimization (HPO) of Neural Networks (NNs) is a computationally expensive procedure. On accelerators, such as NVIDIA Graphics Processing Units (GPUs) equipped with Tensor Cores, it is possible to speed-up the NN training by reducing the precision of some of the NN parameters, also referred to as mixed precision training. This paper investigates the performance of three popular HPO algorithms in terms of the achieved speed-up and model accuracy, utilizing early stopping, Bayesian, and genetic optimization approaches, in combination with mixed precision functionalities. The benchmarks are performed on 64 GPUs in parallel on three datasets: two from the vision and one from the Computational Fluid Dynamics domain. The results show that larger speed-ups can be achieved for mixed compared to full precision HPO if the checkpoint frequency is kept low. In addition to the reduced runtime, small gains in generalization performance on the test set are observed.
AB - Hyperparameter Optimization (HPO) of Neural Networks (NNs) is a computationally expensive procedure. On accelerators, such as NVIDIA Graphics Processing Units (GPUs) equipped with Tensor Cores, it is possible to speed-up the NN training by reducing the precision of some of the NN parameters, also referred to as mixed precision training. This paper investigates the performance of three popular HPO algorithms in terms of the achieved speed-up and model accuracy, utilizing early stopping, Bayesian, and genetic optimization approaches, in combination with mixed precision functionalities. The benchmarks are performed on 64 GPUs in parallel on three datasets: two from the vision and one from the Computational Fluid Dynamics domain. The results show that larger speed-ups can be achieved for mixed compared to full precision HPO if the checkpoint frequency is kept low. In addition to the reduced runtime, small gains in generalization performance on the test set are observed.
KW - High-Performance Computing
KW - Hyperparameter Optimization
KW - Mixed Precision
UR - https://www.scopus.com/pages/publications/85178163613
U2 - 10.1145/3624062.3624259
DO - 10.1145/3624062.3624259
M3 - Conference contribution
T3 - Proceedings of the SC '23 Workshops of The International Conference on High Performance Computing, Network, Storage, and Analysis
SP - 1776
EP - 1779
BT - Proceedings of 2023 SC Workshops of the International Conference on High Performance Computing, Network, Storage, and Analysis, SC Workshops 2023
PB - Association for Computing Machinery, Inc
Y2 - 12 November 2023 through 17 November 2023
ER -