Role Description
Deploy, integrate, and operate high-performance storage for GPU-accelerated compute and AI platforms. You will own the storage layer where Kubernetes meets bare metal β standing up NFS-based high-performance storage, wiring it into clusters via CSI, and tuning it to keep data flowing to GPU workloads at scale. Work spans hybrid, edge, and air-gapped deployments built on the Mirantis K0rdent stack.
We are looking for a senior DevOps engineer who treats storage as infrastructure to be automated, observed, and tuned β not hand-managed. The right candidate is fluent in Kubernetes storage, deeply versed in Linux storage and networking fundamentals down to the kernel and NFS-client layer, and knows how to make high-performance NAS actually perform under demanding workloads. You should reach for infrastructure-as-code and GitOps by default, be self-directed in diagnosing performance and reliability issues end to end, set operational standards for others to follow, and communicate clearly across teams. Bare-metal hardware experience is a strong plus, but deep Linux storage knowledge is essential.
Responsibilities
-
Storage Integration & Operation
-
Integrate NFS-based high-performance storage (e.g., VAST, Dell PowerScale) into Kubernetes clusters via CSI, storage classes, and persistent volumes.
-
Tune the NFS data path β mount options, nconnect/RDMA, Linux client, and network settings β for high-throughput, low-latency GPU/AI workloads.
-
Deploy and operate storage services and operators; manage capacity, quotas, snapshots, and lifecycle.
-
Linux Platform & System Integration
-
Configure and optimize Linux systems for storage workloads, including driver setup, file system layout, network tuning, and kernel parameter optimization.
-
Deliver storage integration for k0s-based Kubernetes via Cluster API (CAPI) and K0rdent management/child cluster topologies.
-
Operate storage in fully disconnected (air-gapped) environments, including local artifact/mirror connectivity (Harbor) and PKI/TLS considerations.
-
Automation & Observability
-
Automate storage provisioning and configuration with infrastructure-as-code (Terraform/OpenTofu) and GitOps pipelines (ArgoCD or Flux).
-
Build monitoring, alerting, and observability for storage performance, capacity, and health.
-
Diagnose and resolve performance, reliability, and scaling issues across the storage stack.
Qualifications
-
7+ years of experience in SRE or infrastructure operations
-
5+ years of building/operating distributed production Storage systems at scale
-
Hands-on with High Performance Storage solutions (VAST, Weka, DDN, PowerScale)
-
Linux and Kubernetes storage fundamentals (NFS, CSI)
Requirements
-
Bare-metal experience: hands-on experience with bare-metal host provisioning, raw disk/hardware layout, and physical server storage configurations.
-
Hands-on experience with VAST and/or Dell PowerScale.
-
Experience with GPUDirect Storage and RDMA/RoCE data paths
-
Experience with the Mirantis K0rdent stack (K0rdent Enterprise, K0rdent AI, k0s, MKE) and Cluster API.
-
Familiarity with other storage backends (Ceph, object/S3) and CSI driver operations.
-
Proven experience in sovereign or high-security air-gapped environments.
Benefits
-
Work with an established Silicon Valley leader in the cloud infrastructure industry;
-
Work with exceptionally passionate, talented and engaging colleagues, helping Fortune 500 and Global 2000 customers implement next-generation cloud technologies;
-
Be a part of cutting-edge, open-source innovation;
-
Thrive in the high-energy environment of a young company where openness, collaboration, risk-taking, and continuous growth are valued;
-
Professional development and training;
-
Attend conferences and working groups;
-
Company outings, happy hours, hackathons, and tech talks;
-
Receive a competitive compensation package with a strong benefits plan.