Role Description
We are seeking an experienced DevOps / Site Reliability Engineer (SRE) to design, automate, and scale our cloud-native infrastructure pipelines. The ideal candidate will bridge the gap between software development and systems operations, building highly available deployment pipelines, writing infrastructure-as-code (IaC), and optimizing cluster scaling to maximize application uptime, system reliability, and performance.
Key Responsibilities
-
Design, build, and optimize automated CI/CD pipelines across cloud platforms using industry-standard automation servers (e.g., GitHub Actions, GitLab CI, Jenkins, ArgoCD).
-
Architect and manage infrastructure-as-code (IaC) templates using Terraform or OpenTofu to provision secure, modular, and repeatable multi-environment architectures.
-
Orchestrate containerized production workloads, configuring cluster scaling, service meshes, network routing policies, and deployment strategies on Kubernetes (EKS/AKS/GKE).
-
Implement automated monitoring, logging, and alerting systems utilizing observability tools (e.g., Prometheus, Grafana, Datadog, ELK stack) to actively track platform performance metrics.
-
Drive system high-availability and fault tolerance efforts, designing disaster recovery plans, automated load balancing parameters, and self-healing cluster scripts.
-
Manage centralized configuration and secret management systems, securely vaulting database credentials, API tokens, and certificate profiles (e.g., HashiCorp Vault, AWS Secrets Manager).
-
Participate in on-call rotations and lead incident root-cause analysis (RCA), systematically diagnosing runtime infrastructure failures, performance bottlenecks, and resource leaks.
Qualifications
-
5 to 9 years of core systems engineering or software development experience, with 4+ dedicated years actively designing, building, and operating cloud-native production platforms.
-
Strong technical mastery of Kubernetes cluster administration, Terraform automation layouts, Linux system internals, shell scripting (Bash, Python, or Go), and network protocols.
-
Deep structural understanding of microservices design, caching mechanics, database scaling limits, and cloud provider API governance.
-
Mandatory certification: AWS DevOps Professional, Azure DevOps Expert, or CKA.
Requirements
-
Prior experience implementing DevSecOps controls (e.g., integrating SAST/DAST tools directly into container build phases).
-
Experience with GitOps methodologies and progressive delivery mechanisms (e.g., Canary or Blue/Green deployments using Flagger or Istio).
Benefits
-
12+ years of total IT software engineering or operational management background, with 6+ dedicated years acting as a CISO, Director of Security, or Principal Enterprise GRC Advisor.
-
Strong visionary mastery of modern security trends, zero-trust target states, risk calculation paradigms, and multi-cloud information landscape parameters.
-
Deep communication execution skills, with a proven history of negotiating security budgets, steering board panels, and handling high-pressure public communication events.
-
Mandatory certification: CISM or CISSP.