Role Description
This role is responsible for designing, building, and supporting scalable internal platform and infrastructure capabilities that enable engineering teams to deliver software efficiently and reliably. The ideal candidate brings strong experience across cloud infrastructure, Kubernetes, CI/CD, GitOps, and observability practices, along with a collaborative mindset and strong communication skills. This individual will help improve developer experience, automate operational processes, and support modern cloud-native engineering practices across a multi-cloud environment.
-
Design, provision, and maintain scalable cloud infrastructure across Azure, AWS, and/or GCP
-
Support Kubernetes-based platforms, including deployment automation, scaling, networking, and observability
-
Build and maintain CI/CD and GitOps workflows using modern automation and infrastructure-as-code (IaC) practices
-
Develop reusable tooling, templates, and platform capabilities to improve developer productivity and consistency
-
Implement and support observability, monitoring, alerting, and reliability practices across platform environments
-
Collaborate with application, security, networking, and infrastructure teams to align platform standards and delivery processes
-
Support platform security, access management, secrets handling, compliance-related controls
-
Troubleshoot infrastructure and deployment issues across distributed environments
-
Contribute to operational documentation, standards, and continuous improvement initiatives
Qualifications
-
5+ years of experience in Platform, DevOps, Site Reliability (SRE), or related Engineering roles
-
Strong hands-on experience with Kubernetes in production environments
-
Experience with infrastructure-as-code tools such as Terraform or similar platforms
-
Familiarity with GitOps and CI/CD tooling (e.g., Flux, ArgoCD, GitHub Actions, Jenkins, Harness)
-
Experience supporting cloud-native environments across Azure, AWS, or GCP
-
Proficiency with scripting or automation languages such as Python, Go, or Bash
-
Experience with observability and monitoring platforms (e.g. Dynatrace, Datadog, Grafana, etc.)
-
Understanding of networking, security, and access control concepts within cloud environments
-
Familiarity with Docker, container registries, Helm, and related ecosystem tooling
-
Strong communication skills and ability to work collaboratively within distributed teams
Preferred Skills & Experience
-
Experience with service mesh technologies or internal developer platforms
-
Familiarity with cloud cost optimization and FinOps concepts
-
Relevant cloud, Kubernetes, or infrastructure certifications
-
Experience supporting large-scale engineering organizations or shared platform environments