Role Description
The Senior Cloud Operations Reliability Engineer is responsible for driving operational excellence and strengthening the reliability posture of cloud-based services and supported platforms. This role owns critical reliability initiatives, establishes observability and service health practices, and leads incident response coordination to improve service availability, resiliency, and recovery.
-
Own service reliability and operational healthβestablish and maintain SLOs/SLIs, design monitoring and alerting strategies, and drive improvements that enhance service availability and performance across cloud platforms.
-
Lead incident response coordination and post-incident processes, including troubleshooting complex production issues, conducting root cause analysis, and driving remediation activities with accountability for timeline and resolution quality.
-
Design and implement reliability-focused automation, operational tooling, and runbooks to reduce manual toil, improve response consistency, and strengthen production readiness and resilience; apply Infrastructure as Code practices where appropriate to support recovery, reliability, and operational consistency.
-
Build observability solutions through comprehensive monitoring, logging, and alerting strategies; establish event correlation and escalation procedures to ensure rapid problem detection and response.
-
Conduct performance and capacity analysis, evaluate utilization trends, identify bottlenecks; provide recommendations for reliability-focused scaling, performance improvement, capacity planning, and operational readiness of cloud-based services.
-
Partner with development and engineering teams to evaluate deployment readiness, support deployment reliability improvements, and implement operational best practices that strengthen service reliability, rollback readiness, and production supportability.
-
Contribute to disaster recovery and business continuity planning, conduct operational readiness exercises, and ensure recovery procedures and documentation reflect current production state and evolving business requirements.
-
Mentor team members and establish reliability standards and practices within Cloud Operations and supported service areas; create and maintain operational documentation, standard operating procedures, and knowledge base materials.
-
Support operational adherence to cloud governance, compliance, and security initiatives; including access control, tagging, logging, and audit readiness, and reliability-related documentation.
-
Perform other duties that support the overall objective of the position.
Qualifications
-
Bachelor's degree in Computer Science, Engineering, Information Systems, or a related field.
-
Any combination of education and experience which would provide the required qualifications for the position.
Requirements
-
10+ years of professional experience in Cloud Operations, Site Reliability Engineering, DevOps, Infrastructure Operations, or a related discipline with demonstrated ownership of production systems.
-
Extensive hands-on experience supporting production cloud environments using Google Cloud Platform (GCP), AWS, or equivalent cloud service providers.
-
Proven expertise in monitoring, observability platforms, alerting strategies, incident response, root cause analysis, and production support in distributed or cloud-native architectures.
-
Demonstrated experience with Infrastructure as Code (Terraform, Deployment Manager, CloudFormation, etc.) and version control best practices.
-
Strong background in incident management and post-incident review processes; experience driving corrective actions and establishing reliability improvements.
-
Experience with Kubernetes operations, containerization, and orchestration platforms.
-
Experience with application performance monitoring (APM) and distributed tracing.
-
Experience mentoring junior engineers or leading operational improvements initiatives.
License/Certification Required
-
Google Cloud certifications: Google Cloud Associate Cloud Engineer, Google Cloud Professional Cloud Architect, Google Cloud Professional Cloud Operations Engineer, or Google Cloud Professional Data Engineer.
-
AWS certification: AWS SysOps Administrator or equivalent.
-
Advanced certifications in Kubernetes, Terraform, observability platforms, DevOps, Site Reliability Engineering (SRE), or ITIL.
Knowledge, Skills & Abilities
-
Knowledge of CI/CD practices, cloud governance, compliance frameworks, disaster recovery, and business continuity planning.
-
Deep technical knowledge of Google Cloud Platform (GCP), AWS, or similar cloud providers; understanding of cloud-native services, networking, security, and compute models.
-
Familiarity with observability tools such as Grafana, Prometheus, Cloud Monitoring, or similar platforms.
-
Security operations, compliance auditing, or audit readiness processes, preferred.
-
Hands-on expertise with monitoring platforms (Datadog, New Relic, Prometheus, Cloud Monitoring, etc.); ability to design effective dashboards, alerts, and health checks.
-
Proficiency in scripting languages (Python, Bash, Go, etc.) to develop automation solutions that reduce manual effort.
-
Advanced ability to diagnose complex, multi-layered infrastructure issues and coordinate timely recovery.
-
Ability to translate complex technical findings into actionable recommendations; experience influencing cross-functional teams on reliability practices.