Role Description
Lead the implementation and continuous improvement of Site Reliability Engineering (SRE) practices including:
-
Service Level Indicators (SLIs)
-
Service Level Objectives (SLOs)
-
Error Budgets to improve system reliability and operational excellence
Design, deploy, and support highly available, scalable, and resilient cloud-native platforms across:
-
AWS
-
Azure
-
Kubernetes environments
-
Build and maintain observability solutions utilizing:
-
Datadog
-
Splunk
-
Grafana
-
OpenTelemetry
-
Prometheus
-
Develop operational dashboards, intelligent alerting, service health scorecards, and reliability metrics to improve incident detection and operational awareness
-
Lead production incident response activities, root cause analysis (RCA), and post-incident remediation efforts to improve service stability and reduce recurrence
-
Develop and maintain Infrastructure as Code (IaC) solutions using Terraform and cloud automation technologies
-
Build automation and self-healing capabilities to reduce operational toil, improve platform resilience, and accelerate incident resolution
-
Partner with software engineering, architecture, security, and platform teams to improve production readiness, reliability, scalability, and performance
-
Support and optimize Kubernetes-based container platforms and microservices running in production environments
-
Design and implement CI/CD and GitOps practices utilizing:
-
GitHub Actions
-
ArgoCD
-
Azure DevOps
-
Drive adoption of AI-enabled operational capabilities including:
-
Anomaly detection
-
Intelligent alerting
-
Incident automation
-
Operational analytics
-
Mentor engineers on SRE principles, observability, automation, operational excellence, and cloud-native best practices
Qualifications
-
Bachelorβs degree in Computer Science, Engineering, Information Technology, or related field
-
7+ years of experience in Site Reliability Engineering, DevOps, Platform Engineering, Cloud Engineering, or Software Engineering
-
4+ years of hands-on experience supporting cloud infrastructure in AWS, Azure, or GCP environments
-
3+ years of hands-on experience managing Kubernetes platforms including EKS, AKS, or GKE in production environments
-
3+ years of Infrastructure as Code (IaC) experience using Terraform or equivalent automation technologies
-
3+ years of hands-on experience with observability and monitoring platforms such as Datadog, Splunk, Dynatrace, Grafana, Prometheus, OpenTelemetry, or similar solutions
-
3+ years of experience implementing and supporting monitoring, logging, distributed tracing, alerting, SLIs, SLOs, and Error Budget frameworks
-
3+ years of experience building and supporting CI/CD pipelines using GitHub Actions, Azure DevOps, Jenkins, ArgoCD, or equivalent technologies
-
3+ years of experience with scripting and automation skills using Python, Bash, PowerShell, or similar languages
-
Ability to participate in rotating on-call support schedules
Preferred Qualifications
-
Experience supporting mission-critical production systems and leading incident response and root cause analysis activities
-
Strong understanding of distributed systems, cloud-native architectures, networking, security, IAM, encryption, and reliability engineering principles
-
Proven ability to collaborate effectively across engineering, platform, architecture, and security teams
-
Experience implementing enterprise observability solutions using Datadog APM, Splunk Observability Cloud, Dynatrace, Grafana, or OpenTelemetry
-
Experience with AIOps, intelligent alerting, anomaly detection, operational automation, and predictive analytics platforms
-
Experience supporting AI/ML, Generative AI, Large Language Models (LLMs), Retrieval-Augmented Generation (RAG), or data-intensive workloads in production environments
-
Experience with GitOps frameworks such as ArgoCD or Flux
-
Experience supporting multi-region and multi-cluster cloud deployments
-
Experience mentoring engineers and leading reliability improvements across multiple teams
-
Experience working within regulated environments such as Healthcare, HIPAA, SOC2, NIST, or FedRAMP
-
Industry certifications such as Certified Kubernetes Administrator (CKA), AWS Solutions Architect, Azure Solutions Architect Expert, HashiCorp Terraform Associate, or equivalent cloud certifications
Benefits
-
Comprehensive benefits package
-
Incentive and recognition programs
-
Equity stock purchase
-
401k contribution (all benefits are subject to eligibility requirements)