Role Description
Senior DevOps Engineer to support growing cloud and platform initiatives. The team needs an experienced resource who can immediately contribute to AWS/Kubernetes-based projects, infrastructure automation, CI/CD, observability, and AI-enabled engineering without extensive training. This is additional project support and capacity expansion, not a simple backfill.
Client is seeking a Senior DevOps Engineer to join their team on a remote-friendly contract basis. This is a hands-on role for an engineer with strong experience in AWS, Kubernetes, Terraform, and CI/CD who can independently drive reliable infrastructure decisions in ambiguous environments. You will be instrumental in:
-
Automating and operating our cloud and Kubernetes platform
-
Improving deployment flows
-
Strengthening security and observability
-
Ensuring systems are scalable, reliable, and maintainable
-
Using modern AI-assisted engineering workflows to accelerate analysis, troubleshooting, documentation, and delivery while maintaining high standards for correctness, security, and operational safety
Responsibilities
-
Design, implement, and operate cloud infrastructure and automation using AWS and Terraform across services such as EC2, S3, EKS, ECR, Route 53, IAM, and core networking components.
-
Own EKS and Kubernetes platform operations, including cluster lifecycle management, upgrades, node group strategy, autoscaling, workload scheduling, capacity planning, performance tuning, and cost optimization.
-
Build and improve core Kubernetes platform capabilities such as networking, ingress, DNS integration, service discovery, RBAC, policy enforcement, secrets management, and multi-environment consistency.
-
Develop and maintain deployment workflows using Helm, Kustomize, GitLab CI/CD, and GitOps approaches with tools such as Argo CD or Flux to improve release speed, consistency, traceability, and rollback safety.
-
Deploy and evolve observability capabilities across infrastructure and Kubernetes environments, including metrics, logging, alerting, dashboards, and incident diagnostics.
-
Support and optimize platform and application workloads on Kubernetes, improving deployment patterns, scaling behavior, runtime efficiency, resilience, and day-to-day operational support.
-
Strengthen security across infrastructure, pipelines, and Kubernetes environments through IAM least-privilege access, secrets management, image and artifact scanning, admission controls, policy-as-code, workload isolation, and practical runtime hardening.
-
Support the reliability, backup integrity, availability, and operational performance of PostgreSQL/RDS environments in partnership with application teams.
-
Work closely with development and platform teams to improve deployment strategies, runtime reliability, developer experience, and operational standards for cloud-native systems.
-
Monitor system health, investigate incidents, troubleshoot infrastructure and application issues, and drive timely resolution through strong root cause analysis and preventative improvements.
-
Lead technical decision-making within the scope of the role by prioritizing work, evaluating tradeoffs, integrating inputs from multiple stakeholders, and driving issues through to completion with limited oversight.
-
Use modern AI tools to accelerate infrastructure design exploration, Terraform authoring, Kubernetes troubleshooting, CI/CD workflow drafting, observability analysis, and operational documentation, while rigorously validating outputs for correctness, security, maintainability, and production readiness.
-
Document and continually refine DevOps methodologies, infrastructure standards, deployment workflows, operational procedures, and support runbooks.
Qualifications
-
7+ years of experience in DevOps, platform engineering, SRE, or related roles, with a proven track record in designing and operating scalable cloud infrastructure.
-
Strong expertise with AWS services, including EC2, S3, EKS, ECR, Route 53, IAM, and core networking and security concepts.
-
Advanced hands-on experience with Kubernetes, including cluster operations, upgrades, autoscaling, networking, ingress, RBAC, policy enforcement, and production workload management.
-
Strong experience with Terraform for infrastructure-as-code, environment management, and repeatable platform provisioning.
-
Experience with Kubernetes packaging and deployment tools such as Helm and Kustomize.
-
Experience with CI/CD systems such as GitLab CI/CD and familiarity with GitOps approaches using Argo CD, Flux, or similar tools.
-
Strong experience with observability and monitoring tools such as Prometheus, Grafana, Loki, or comparable tooling.
-
Experience evaluating or operating Kubernetes ecosystem tools such as ingress controllers, operators, Cluster Autoscaler, or Karpenter.
-
Strong experience with PostgreSQL/RDS, including performance tuning, reliability, and operational support.
-
Solid understanding of cloud and container security practices, including IAM least privilege, secrets management, image scanning, admission controls, and policy enforcement.
-
Strong scripting and automation skills in Python, Bash, or similar languages.
-
Practical fluency with AI-assisted engineering tools such as LLMs, code assistants, and workflow automation tools, along with strong judgment to evaluate and refine AI-generated outputs responsibly.
-
Excellent problem-solving, teamwork, and communication skills.
-
Ability to work independently in a remote-friendly environment and make sound technical decisions in ambiguous situations.
Requirements
-
7+ years of relevant DevOps/Cloud Infrastructure experience
-
Strong production support and deployment experience
-
Comfortable owning projects independently
Education
-
Bachelor's Degree (preferred) or higher in Computer Science, Information Technology, or a related field, or equivalent practical experience.
-
Strong practical experience can fully compensate for lack of degree.
Top Required Skills (Must-Have)
-
AWS Cloud Engineering (strong hands-on expertise)
-
Kubernetes Administration & Deployment
-
Terraform / Infrastructure as Code
-
Docker
-
Networking & Security Fundamentals
-
Observability & Monitoring
-
Senior-level DevOps experience
Desirable Skills
-
Experience with additional cloud platforms such as GCP or Azure.
-
Experience improving incident response, operational readiness, and platform standards for distributed engineering teams.
-
Experience designing internal platform capabilities that improve developer self-service and reduce operational toil.
-
Familiarity with compliance, governance, and risk management requirements in regulated or enterprise environments.
-
AI-assisted Engineering
Certifications Preferred
-
AWS Certifications
-
Kubernetes Certifications (CKA/CKAD preferred)
-
Terraform Certification
Interview Process
-
Manager prefers one interview: Teams interview with manager and additional team members.
-
Expected to start with one interview round.
-
May expand based on hiring manager preference.
-
Technical deep-dive expected.
-
Candidates should be prepared for questions exposing strengths and weaknesses in:
-
AWS
-
Kubernetes
-
Terraform
-
DevOps concepts
-
AI/LLM experience
-
Looking for authentic expertise, not buzzword-heavy answers.