Role Description
We are seeking a Lead Site Reliability Engineer to strengthen critical infrastructure reliability and accelerate DevOps maturity for high-impact services. You will design scalable automation, improve CI/CD and release practices, and lead rapid incident response.
-
Design reliability strategies and SRE practices for business-critical infrastructure
-
Build automation and tooling in Python to improve stability, consistency, and operational leverage
-
Develop and maintain CI/CD workflows and source control practices using GitLab
-
Lead incident response during on-call rotations and restore service for business-critical issues
-
Improve release management processes to support enterprise-scale delivery
-
Harden cloud infrastructure across networking, compute, security, and IAM controls
-
Implement configuration automation to reduce manual work and prevent drift
-
Operate and troubleshoot Kubernetes-based workloads and developer-facing platform usage
-
Partner with engineering stakeholders to prioritize reliability work and manage change safely
-
Assess systemic risks and drive corrective actions to prevent recurring incidents
Qualifications
-
5+ years of site reliability engineering or DevOps experience in cloud environments
-
Hands-on experience with a leading cloud provider, with practical work across Amazon Web Services and Microsoft Azure
-
Leadership ability to guide technical direction and take ownership of critical infrastructure outcomes
-
Project delivery experience improving DevOps tools, processes, and engineering maturity at scale
-
Deep CI/CD knowledge across pipelines, source control, and release management
-
Strong Kubernetes skills with practical usage as a developer
-
Advanced Python programming skills for automation and tooling
-
Enterprise-scale release management experience supporting complex systems
-
Solid infrastructure fundamentals across networking, compute, security, IAM, and configuration automation
-
Strong analytical skills to diagnose complex issues and identify high-leverage solutions
-
Effective incident response skills, including on-call ownership and rapid restoration of service
-
Upper-Intermediate English proficiency (B2, Upper-Intermediate)
Requirements
-
Nice to have Amazon Web Services certification or proven advanced AWS operational experience
-
Microsoft Azure certification or proven advanced Azure operational experience
-
AI Architecture experience for reliability-focused platform design
-
AI Solution Engineering experience integrating AI-enabled capabilities into operations
-
Gen AI Solutions Development experience for operational intelligence and automation use cases