Role Description
We are seeking an experienced Senior Cloud Infrastructure & AIOps Engineer to join our platform engineering team. This role is critical to building and maintaining our cloud infrastructure while leveraging AI-driven operational tools to automate, optimize, and scale our systems. The ideal candidate will have deep expertise in cloud services, infrastructure-as-code, identity management, and emerging AIOps technologies.
Key Responsibilities
-
Design, build, and maintain scalable cloud infrastructure on AWS and/or GCP using vendor-agnostic approaches and best practices.
-
Develop and maintain Infrastructure-as-Code using Terraform for reproducible, version-controlled infrastructure deployments.
-
Implement and manage Identity & Access Management (IAM) solutions, including Okta and Auth0 integrations for secure authentication and authorization.
-
Lead AIOps initiatives using AI-driven operational tools to automate monitoring, incident response, and remediation processes.
-
Mentor junior engineers on cloud infrastructure best practices, DevOps fundamentals, and automation techniques.
-
Collaborate with development and operations teams to ensure infrastructure reliability, security, and performance.
-
Implement CI/CD pipelines and deployment automation to enable rapid, safe releases.
-
Monitor, troubleshoot, and optimize cloud infrastructure performance, costs, and security posture.
-
Participate in on-call rotations and incident response to ensure high availability and quick recovery.
Qualifications
-
6-9 years of experience in DevOps, cloud infrastructure, or site reliability engineering.
-
Strong hands-on experience with AWS and/or GCP cloud platforms, with understanding of multi-cloud and vendor-agnostic architecture patterns.
-
Expert-level proficiency with Terraform for infrastructure provisioning and management.
-
Demonstrated experience implementing and managing identity and access management solutions (Okta and/or Auth0).
-
Solid understanding of DevOps fundamentals including containerization, orchestration, and automation.
-
Experience with AIOps concepts and tools for operational intelligence, predictive analytics, and automated remediation.
-
Proficiency in scripting languages (Python, Bash, Go, or similar) for automation and tooling.
-
Strong knowledge of containerization (Docker) and container orchestration platforms (Kubernetes).
-
Excellent communication and problem-solving skills with the ability to mentor and lead engineering teams.
-
Relevant cloud certifications (AWS, GCP, or similar).
Preferred Qualifications
-
Experience with machine learning operations (MLOps) and AI/ML infrastructure deployment.
-
Hands-on experience with advanced monitoring and observability platforms (Dynatrace, New Relic, Prometheus, ELK stack).
-
Knowledge of security best practices, including infrastructure security hardening and vulnerability management.
-
Track record of implementing incident management and disaster recovery strategies.
-
Contributions to open-source infrastructure or DevOps projects.
What We're Looking For
-
Above all, we're seeking a technically excellent engineer who is passionate about AIOps and leveraging AI to solve operational challenges and improve system reliability.
-
Thinks strategically about cloud architecture while maintaining pragmatism about trade-offs and costs.
-
Values automation, repeatability, and treating infrastructure as code.
-
Is committed to continuous learning and staying current with rapidly evolving cloud and AI technologies.
-
Excels at collaboration and enjoys working across teams to solve complex technical problems.
-
Demonstrates ownership and takes pride in building reliable, secure, and scalable systems.