Role Description
Weβre looking for a DevOps Engineer who will help build, maintain, and evolve the infrastructure behind our products and engineering teams. You will be responsible for designing and managing cloud infrastructure, Kubernetes environments, CI/CD pipelines, observability, reliability, and production operations across multiple products. Youβll work closely with Backend, Product, and Engineering teams to ensure our systems remain scalable, stable, secure, and efficient as traffic, workloads, and product complexity continue to grow.
-
Design, implement, and maintain scalable cloud infrastructure.
-
Establish and manage CI/CD pipelines for streamlined and reliable software deployment.
-
Manage and maintain services running on Kubernetes clusters, troubleshoot failures, and optimise performance.
-
Troubleshoot and resolve performance, availability, and reliability issues across distributed systems.
-
Handle production incidents, conduct post-mortems, identify root causes, and improve reliability processes.
-
Implement and continuously improve observability practices, including logging, monitoring, tracing, and alerting.
-
Monitor infrastructure resource usage and optimise system performance and cloud efficiency.
-
Manage global CDN configurations to ensure stable and efficient content delivery.
-
Maintain SSO systems, databases, backups, and recovery processes.
-
Provide technical support for development, staging, and production environments.
-
Automate infrastructure and operational processes using scripts and Infrastructure as Code.
Qualifications
-
Strong hands-on experience with AWS and a solid understanding of cloud architecture and operations.
-
Strong knowledge of Kubernetes and Docker, including experience managing production-grade clusters.
-
Experience with managed Kubernetes platforms such as EKS, GKE, or AKS.
-
Advanced experience with Terraform, Helm, and Git.
-
Hands-on experience implementing and maintaining GitHub Actions for CI/CD pipelines.
-
Experience with monitoring and observability tools such as Prometheus, Grafana, and OpenTelemetry.
-
Ability to automate system and deployment tasks using Python and Bash.
-
Strong troubleshooting skills and experience working with distributed production systems.
-
Experience using AI-assisted development tools such as Claude, Cursor, GitHub Copilot, ChatGPT, Codex, or similar.
Requirements
-
Hands-on experience managing API gateways for high-traffic applications (Nice to have).
-
Experience working with machine learning projects, including deploying, monitoring, and scaling ML workloads (Nice to have).
-
Familiarity with machine learning platforms, frameworks, or libraries (Nice to have).
-
Experience in a DBA role, including managing, maintaining, and optimising large-scale databases (Nice to have).
Benefits
-
Innovative Environment: We're all about trying new things and pushing the envelope in EdTech.
-
Global remote work.
-
Health: company-provided medical expense compensation.
-
AI solutions: AI subscription and other tools.
-
Balance: Flexible paid time off, you get 21 days of annual leave + 10 bank holidays.
-
Collaborative Culture: Work alongside passionate professionals who are as driven as you are.
Interview Stages
-
HR Interview
-
Technical Interview
-
Final Interview
Join StellarTech!