Role Description
The role requires extensive experience in Platform Engineering, Cloud Engineering, SRE, DevOps, or Infrastructure Engineering.
-
10 years of hands-on experience in relevant roles.
-
Priority 1 β Must Have (Hands-On):
-
Azure Cloud Platform:
-
Strong hands-on experience supporting Azure workloads in production environments.
-
Experience designing, building, and supporting Azure infrastructure using Terraform.
-
Good understanding of Azure networking and connectivity patterns.
-
Experience supporting:
-
AKS
-
Application Gateway
-
Azure Traffic Manager
-
Key Vault
-
Azure Monitor / Log Analytics
-
Managed Identities
-
Service Principals
-
Private Endpoints
-
VNets, NSGs, and Route Tables
-
Kubernetes / AKS:
-
Strong practical experience operating and supporting AKS.
-
Ability to troubleshoot:
-
Pod failures
-
Ingress issues
-
DNS issues
-
SSL/TLS certificate issues
-
Network routing issues
-
Performance and availability incidents
-
Experience with:
-
Helm
-
Ingress Controllers
-
Cluster upgrades
-
Scaling
-
Monitoring
-
Terraform:
-
Strong hands-on experience writing and maintaining Terraform.
-
Experience creating reusable modules.
-
Experience managing:
-
State files
-
Remote backends
-
Environment promotion
-
Infrastructure lifecycle
-
Linux & Scripting:
-
Strong Linux administration fundamentals.
-
Practical experience troubleshooting production issues.
-
Bash scripting mandatory.
-
Python desirable.
-
Application Support / Troubleshooting:
-
Must be comfortable supporting business applications end-to-end.
-
Expected capability workflow:
-
User reports application unavailable
-
Understand architecture
-
Trace traffic flow
-
Review Application Gateway
-
Review DNS / Networking
-
Review AKS health
-
Review certificates
-
Review monitoring & logs
-
Identify root cause
-
Priority 2 β Highly Desirable:
-
GitHub & DevOps Platform:
-
Hands-on experience with:
-
GitHub Enterprise
-
GitHub Actions
-
Shared workflows
-
Reusable pipelines
-
Repository onboarding
-
Branch protections
-
GitHub security features
-
Experience supporting:
-
Runner issues
-
Disk space issues
-
Network connectivity issues
-
Dependency failures
-
Self-hosted runners lifecycle management
-
API Management
-
Pipeline failures
-
GitOps:
-
Experience with:
-
ArgoCD
-
GitOps deployment models
-
Kubernetes deployment automation
-
Monitoring & Observability:
-
Experience working with:
-
Prometheus
-
Grafana
-
Azure Monitor
-
Log Analytics
-
Application Insights
Qualifications
-
10 years of hands-on experience in Platform Engineering, Cloud Engineering, SRE, DevOps, or Infrastructure Engineering roles.
Requirements
-
Strong hands-on experience supporting Azure workloads in production environments.
-
Experience designing, building, and supporting Azure infrastructure using Terraform.
-
Good understanding of Azure networking and connectivity patterns.
-
Strong practical experience operating and supporting AKS.
-
Strong hands-on experience writing and maintaining Terraform.
-
Strong Linux administration fundamentals.
-
Must be comfortable supporting business applications end-to-end.
Company Description