Role Description
We are seeking a Senior HPC DevOps Engineer to join our Science, Innovation & Labs team, responsible for the scaling, reliability, and automation of our high-performance computing (HPC) and machine learning operations (MLOps) platform.
-
Guide scientists and data teams in navigating and utilizing the platform user interface (UI) effectively, helping them run self-service workloads without direct infrastructure friction.
-
Advise users and manage infrastructure capacity regarding capacity blocks versus on-demand usage, optimizing cost, quotas, and resource availability for heavy workloads.
-
Maintain automated pipelines for infrastructure provisioning and platform service deployments.
-
Resolve technical queries regarding job scheduling failures, cluster bottlenecks, and resource quotas.
-
Collaborate with developer experience teams to improve documentation.
-
Collaborate with engineering teams to monitor GPU utilization via tools such as CloudWatch or Prometheus.
-
Manage AWS GPU instance families and allocate block compute for large-scale ML training and inference pipelines.
-
Ensure compute availability through capacity planning and reservation management.
-
Deploy containerized environments tuned for HPC and GPU pass-through.
-
Deploy and scale HPC workloads on cloud infrastructure utilizing parallel storage and networking solutions.
Qualifications
-
5+ years of experience in HPC or DevOps engineering roles.
-
Knowledge of MPI, OpenMP, and multi-node GPU communication protocols such as NCCL and GPUDirect.
-
Proven experience managing AWS GPU instance families, including P-series, G-series, and Tranium/Inferentia.
-
Hands-on mastery of AWS Capacity Blocks for ML, On-Demand Capacity Reservations (ODCRs), and Service Quota management.
-
Experience in deployment of containerized environments using Apptainer/Singularity, Docker, or Enroot.
-
Understanding of I/O performance bottlenecks when interfacing with distributed file systems such as Lustre, GPFS, BeeGFS, or AWS FSx for Lustre.
-
Hands-on skill in profiling applications using NVIDIA Nsight or similar tools to locate memory and compute bottlenecks.
-
Experience deploying or scaling HPC workloads on cloud infrastructure utilizing EFA, ParallelCluster, and parallel storage (FSx for Lustre).
-
Proficiency in English at a B2+ level.