Role Description
We are looking for a Senior DevOps Engineer, Reliability & Platform Operations to improve the reliability, operational stability, risk posture, and delivery effectiveness of our cloud infrastructure, Kubernetes platforms, CI/CD workflows, databases, middleware, and production systems.
In this role, you will work closely with:
-
DevOps Engineers / Software Delivery Engineers
-
Software Engineering
-
Security
-
Database providers
-
Support
-
Product
-
Leadership teams
Your responsibilities will include:
-
Improve reliability engineering practices across production systems, including SLIs, SLOs, error budgets, postmortems, runbooks, service health indicators, and corrective-action tracking.
-
Lead reliability improvements across cloud platforms, AKS, CI/CD pipelines, application environments, databases, and supporting middleware.
-
Operate and improve AKS clusters, including upgrades, autoscaling, node pools, networking, storage, identity, security, reliability, and operational standards.
-
Provide guidance, best practices, and standards for Kubernetes application deployment patterns.
-
Improve delivery reliability through advanced GitHub Actions workflows, reusable automation, secure deployment patterns, rollback support, and pipeline optimization.
-
Build automation, self-service capabilities, reusable infrastructure patterns, and operational procedures to reduce toil, ticket-driven work, manual handoffs, and infrastructure drift.
-
Define, implement, and maintain Infrastructure as Code and configuration management practices using approved tools.
-
Own and maintain Ansible-based configuration management for OS, middleware, application-supporting services, and operational automation.
-
Support and improve Azure cloud environments as the primary platform, with AWS and GCP experience preferred.
-
Implement and improve observability across metrics, logs, traces, dashboards, alerting, APM, capacity planning, performance tuning, and incident dashboards.
-
Support incident response for production, deployment, infrastructure, database, middleware, and application reliability issues.
-
Operate with production awareness and ownership, supporting reliability improvements and production-risk reduction.
-
Recommend blocking, delaying, or rolling back releases when reliability, security, operational, or business-continuity risks are identified.
-
Understand PHP Symfony monolith behavior sufficiently to support production reliability and collaborate effectively with software engineering teams.
-
Operate, tune, monitor, back up, troubleshoot, and support MySQL, ProxySQL, and RabbitMQ directly.
-
Strengthen security and compliance practices across infrastructure, CI/CD, and runtime environments.
-
Improve backup, restore, disaster recovery, cloud cost visibility, cost control, and operational resilience across critical systems.
-
Mentor DevOps Engineers / Software Delivery Engineers through technical guidance, design review, standards definition, operational knowledge sharing, and reliability best practices.
-
Produce and maintain runbooks, troubleshooting guides, deployment procedures, reliability standards, architecture notes, and knowledge-base articles.
-
Propose technical initiatives related to reliability, platform operations, observability, automation, incident reduction, cloud maturity, and operational risk reduction.
Qualifications
-
Bachelorβs degree in Computer Science, Engineering, Information Technology, or equivalent professional experience.
-
8+ years of experience in DevOps, Site Reliability Engineering, Platform Engineering, cloud operations, infrastructure automation, production operations, or related roles.
-
Strong written and verbal English communication skills.
-
Strong experience operating production systems where reliability, observability, incident response, automation, and operational risk reduction are core responsibilities.
-
Strong hands-on experience with Azure and production Kubernetes / AKS environments.
-
Experience designing or improving CI/CD, Infrastructure as Code, configuration management, observability, and deployment automation at scale.
-
Strong production troubleshooting skills across cloud infrastructure, Linux, containers, Kubernetes, databases, middleware, CI/CD pipelines, and application-supporting services.
-
Ability to understand application behavior and support production reliability of PHP Symfony applications.
-
Experience operating or supporting MySQL, ProxySQL, and RabbitMQ in production environments.
-
Strong understanding of secure delivery and infrastructure security practices.
-
Experience with backup, restore, disaster recovery, business continuity, capacity planning, performance tuning, and cost optimization.
-
Ability to mentor engineers, define technical standards, challenge weak practices constructively, and lead technical initiatives without formal people-management authority.
-
Fully remote work capability, including disciplined written communication, self-management, asynchronous collaboration, and reliable participation in distributed-team workflows.
Requirements
-
Cloud platforms: Azure required; AWS and GCP preferred.
-
Kubernetes: AKS, cluster operations, troubleshooting, upgrades, autoscaling, networking, storage, identity, security, and platform standards.
-
CI/CD: GitHub Actions, reusable workflows, composite actions, environments, approvals, OIDC, self-hosted runners, rollback workflows, artifacts, caching, and pipeline security controls.
-
Infrastructure as Code: Terraform or Pulumi; Pulumi with Python preferred.
-
Configuration management: Ansible.
-
Scripting and automation: Python required; Bash preferred.
-
Containers: Docker / OCI and standalone Docker hosts.
-
Operating systems: Linux administration and troubleshooting.
-
Databases and middleware: MySQL, ProxySQL, RabbitMQ.
-
Observability: New Relic, Logz.io, Azure Metrics, ClickStack preferred.
-
Reliability engineering: SLIs, SLOs, error budgets, postmortems, runbooks, toil reduction, incident response, and corrective actions.
-
Security: HashiCorp Vault, 1Password, secrets management, least privilege, access reviews, dependency scanning, container image scanning, signing/provenance, and CI/CD security controls.
-
Application reliability: PHP Symfony applications production-support context.
-
Operational resilience: backup, restore, disaster recovery, business continuity, capacity planning, and performance tuning.
-
Documentation: runbooks, troubleshooting guides, operational procedures, standards, and knowledge-base articles.
-
Working style: structured problem-solving, strong ownership, production awareness, mentoring, automation mindset, and proactive risk reduction.
Benefits
-
100% remote work from anywhere in Mexico.
-
Competitive monthly salary after taxes.
-
Major Medical Insurance and healthcare coverage.
-
Home office and ergonomics support, including internet and electricity.
-
Professional development opportunities, including English classes.
-
Wellness benefits such as TotalPass gym discounts.
-
Savings plan.
-
Paid time off, including personal days.
-
Collaborative, international, and growth-oriented environment.
-
Opportunity to work with modern cloud, automation, CI/CD, Kubernetes, reliability, observability, and platform technologies.
-
Collaborative engineering culture focused on automation, reliability, operational excellence, and continuous improvement.