Role Description
We are seeking an Azure DevOps Engineer with strong Site Reliability Engineering (SRE) experience and a solid background in software development. This role blends cloud automation, reliability engineering, CI/CD excellence, and hands-on coding to ensure scalable, secure, and highly available systems across Azure environments.
-
Design, build, and maintain Azure DevOps pipelines (YAML-based CI/CD) for application and infrastructure deployments.
-
Implement SRE best practices including observability, incident response, postmortems, error budgets, and service-level objectives (SLOs).
-
Develop automation tools, scripts, and services using C#, Python, PowerShell, or similar languages.
-
Build and maintain IaC using Terraform, Bicep, or ARM templates.
-
Manage and optimize Azure resources including AKS, App Services, Functions, Storage, Key Vault, API Management, and Azure SQL.
-
Implement monitoring and alerting using Azure Monitor, Application Insights, Log Analytics, and Grafana/Prometheus (if applicable).
-
Improve system reliability through performance tuning, capacity planning, chaos engineering, and resilience testing.
-
Collaborate with development teams to ensure applications are designed for reliability, scalability, and cloud-native operation.
-
Support containerization and orchestration using Docker and Kubernetes (AKS).
-
Drive continuous improvement across deployment processes, release automation, and cloud governance.
-
Participate in on-call rotations and lead incident response efforts when needed.
-
Must be an open and apt communicator.
-
Must adapt easily to being embedded in a fast-paced development team.
Qualifications
-
5+ years in Azure DevOps, SRE, or cloud engineering roles.
-
Strong development experience in C#, Python, or PowerShell (must be able to build tools, not just scripts).
-
Hands-on experience with Azure DevOps Repos, Pipelines, Artifacts, and Boards.
-
Deep knowledge of Azure services, cloud architecture, and distributed systems.
-
Strong experience with Kubernetes (AKS), containerization, and microservices.
-
Proficiency with Terraform, Bicep, or ARM templates for IaC.
-
Experience implementing monitoring, logging, tracing, and alerting frameworks.
-
Solid understanding of SRE principles, SLIs/SLOs, error budgets, and reliability metrics.
-
Experience with Git, branching strategies, and DevOps best practices.
Preferred Qualifications
-
Experience with Azure Landing Zones, cloud governance, and enterprise-scale architecture.
-
Familiarity with GitHub Actions, Jenkins, or other CI/CD tools.
-
Experience with chaos engineering tools (Gremlin, Chaos Mesh).
-
Background supporting high-availability, mission-critical systems.
-
Knowledge of security best practices: RBAC, managed identities, secrets management, and zero-trust principles.