Role Description
We are seeking a Senior DevOps Engineer to design, automate, and support our AWS cloud infrastructure and Kubernetes-based application platform. This role will focus on reliability, scalability, security, observability, and deployment automation across production environments.
Qualifications
-
5+ years of experience in DevOps, Infrastructure Engineering, Site Reliability Engineering, Cloud Engineering, or a similar role.
-
Strong hands-on experience with AWS cloud infrastructure.
-
Production experience with Kubernetes, preferably Amazon EKS.
-
Experience with Infrastructure as Code using OpenTofu, Terraform, or similar tools.
-
Experience with Argo CD, Helm, and GitOps-based deployment workflows.
-
Strong scripting or development experience with Python.
-
Experience supporting Aurora PostgreSQL, ElastiCache Redis, and Amazon MQ for RabbitMQ in production environments.
-
Experience with Datadog or similar observability platforms.
-
Strong understanding of cloud networking, including VPCs, subnets, routing, security groups, DNS, TLS, ingress, and load balancing.
-
Comfortable participating in an on-call rotation and supporting production incident response.
-
Strong troubleshooting, communication, and collaboration skills.
Requirements
-
AWS, Amazon EKS, Aurora PostgreSQL, ElastiCache Redis, Amazon MQ for RabbitMQ, AWS Load Balancers, IAM, VPC networking.
-
OpenTofu, Terraform, Infrastructure as Code, Kubernetes, Helm, containers, ingress, services, ConfigMaps, Secrets, autoscaling.
-
Argo CD, GitOps, CI/CD, release automation, Python, Bash or shell scripting, Datadog, metrics, logs, tracing, dashboards, alerting.
-
Experience supporting high-availability SaaS or customer-facing platforms.
-
Experience with Kubernetes autoscaling, ingress controllers, cluster upgrades, and resource optimization.
-
Experience with backup, restore, disaster recovery, and capacity planning.
-
Familiarity with AWS Well-Architected principles and cloud security best practices.
-
Experience with GitHub Actions, GitLab CI, Jenkins, or similar CI/CD tools.
-
Experience supporting message broker platforms such as RabbitMQ or Amazon MQ.
-
AWS or Kubernetes certifications are a plus.
-
Strong ownership mindset, operational discipline, and a passion for automation, reliability, and continuous improvement.
-
Ability to communicate clearly with technical and non-technical stakeholders, with a willingness to mentor others.
Benefits
-
Build and maintain AWS infrastructure using OpenTofu.
-
Operate Kubernetes workloads on Amazon EKS.
-
Manage GitOps deployments using Argo CD and Helm.
-
Develop automation and operational tooling using Python and scripting languages.
-
Support AWS services including Amazon EKS, Aurora PostgreSQL, ElastiCache Redis, Amazon MQ for RabbitMQ, load balancers, IAM, VPC networking, DNS, and security groups.
-
Implement monitoring, dashboards, alerts, logs, metrics, and tracing using Datadog.
-
Troubleshoot production issues and support incident response and root cause analysis.
-
Participate in an on-call rotation to support production systems, respond to incidents, and assist with after-hours maintenance or escalations as needed.
-
Improve CI/CD pipelines, release automation, and deployment reliability.
-
Partner with engineering and security teams to improve infrastructure standards, operational readiness, and cloud security.
-
Create and maintain technical documentation, runbooks, and operational procedures.
Interview Process
-
Interview #1: Video Screen with Talent Acquisition Team
-
Interview #2: Video interview with the Hiring Manager (via MS Teams)
-
Interview #3: Video interview with the Team (via MS Teams)