Senior Site Reliability Engineer @Climavision
All Others
Salary $130,000 - $170..
Remote Location
Employment Type full-time
Posted 1wk ago

[Hiring] Senior Site Reliability Engineer @Climavision

1wk ago - Climavision is hiring a remote Senior Site Reliability Engineer. πŸ’Έ Salary: $130,000 - $170,000 annually πŸ“Location: Australia

Role Description

Are you an experienced Site Reliability Engineer who thrives at the intersection of software engineering and production operations? Do you take pride in keeping mission-critical customer systems reliable under real-world operational pressure? Are you looking for an opportunity to own production reliability for a modern hybrid infrastructure platform spanning cloud, colocation, and edge environments?

Climavision is seeking a Senior Site Reliability Engineer to contribute towards reliability, operational excellence, and production resilience across the company's platform and data services. This role sits on a shared SRE team that supports the full business rather than a single product line, covering both the radar network and the weather intelligence sides of the company as priorities shift. A central focus of this role is building the observability layer that puts the health of the full fleet in one place, and then automating recovery so that systems heal themselves. Multi-cluster and multi-replica high availability across our distributed edge fleet remains a core part of the work.

This is a hands-on engineering role for someone who is equally comfortable troubleshooting Kubernetes clusters, leading incident response, and improving operational maturity across the organization. The successful candidate will combine deep production operations expertise with a disciplined approach to reliability engineering and strong automation skills.

Climavision operates a hybrid infrastructure footprint spanning Microsoft Azure, colocation data centers, and edge Kubernetes clusters, deployed alongside weather radar systems. This role will drive production reliability across Azure, colocation, and edge environments. Right-sizing cluster resources and migrating workloads off Azure to reduce spend are active priorities for the team.

Primary Responsibilities

  • Own production reliability for Climavision's customer-facing platform and data services across Azure, colocation, and edge Kubernetes environments.
  • Work as part of a shared SRE function supporting the whole company rather than a single product line, taking on work across both the radar network and the weather intelligence sides of the business as priorities shift.
  • Contribute to the definition and improvement of SLIs, SLOs, alerting standards, and operational metrics used to measure platform reliability.
  • Build and own the observability layer for the fleet, surfacing data in shared dashboards and building alerting.
  • Design and build automated recovery and self-healing for production systems.
  • Optimize cluster resourcing and cost, including right-sizing workloads and nodes.
  • Support and coordinate production incident response efforts, including troubleshooting and postmortem analysis.
  • Diagnose and resolve complex production issues across application services, Kubernetes infrastructure, storage, and distributed systems.
  • Drive multi-replica and multi-cluster high availability across Climavision's services.
  • Contribute to the multi-cluster high-availability strategy across Climavision's hybrid fleet.
  • Operate and improve Climavision's self-managed Kubernetes platform.
  • Ensure Kubernetes platform lifecycle activities are executed to preserve service availability.
  • Improve reliability and operational maturity of production platform services.
  • Design and validate Kubernetes workloads for resiliency and operational efficiency.
  • Partner with software engineering teams to improve production readiness.
  • Maintain and improve deployment pipelines, Helm charts, and infrastructure automation.
  • Support and evolve Climavision's observability platform.
  • Conduct performance engineering and capacity-planning efforts for customer-facing services.
  • Help facilitate blameless postmortem reviews and drive operational follow-up items.
  • Improve disaster recovery, failover, and business continuity capabilities.
  • Drive operational excellence initiatives.
  • Contribute as a senior technical resource and mentor on reliability engineering practices.

On-Call Expectation

  • Participate in a rotating on-call schedule made up of two separate rotations: a weekday rotation and a weekend rotation.
  • Expect a weekday shift roughly every five weeks and a weekend shift roughly every five weeks.

Qualifications

  • A bachelor's degree in computer science, software engineering, or a related field; equivalent professional experience considered.
  • Minimum of 7 years of experience in Site Reliability Engineering, DevOps, or a related infrastructure-focused role.
  • Deep, hands-on experience operating native Kubernetes.
  • Demonstrated experience optimizing Kubernetes clusters.
  • Experience increasing operational visibility.
  • Experience designing and operating workloads for safe horizontal scaling.
  • Experience designing or operating multi-cluster high-availability architectures.
  • Experience supporting customer-facing production systems.
  • Experience diagnosing and resolving production incidents.
  • Experience operating Kubernetes outside of strictly managed cloud environments.
  • Experience with Kubernetes operational tooling and ecosystem technologies.
  • Strong understanding of infrastructure automation and Infrastructure as Code concepts.
  • Experience supporting CI/CD and production deployment pipelines.
  • Experience with monitoring, logging, and observability platforms.
  • Experience operating distributed systems and microservice-based architectures.
  • Working knowledge of Microsoft Azure infrastructure.
  • Strong troubleshooting skills across infrastructure, application, and platform layers.
  • Demonstrated experience participating in a structured production on-call rotation.
  • Working familiarity with Jira, Confluence, and Microsoft Entra.
  • Strong written and verbal communication skills.
  • Experience working in fast-moving engineering environments.

Nice to have, but not required

  • Experience operating Kubernetes platforms using RKE2 and Rancher.
  • Experience with Octopus Deploy.
  • Experience with SOC 2 or comparable security auditing and compliance work.
  • Experience supporting hybrid cloud and colocation infrastructure environments.
  • Experience with service mesh technologies such as Istio.
  • Experience with Kubernetes-native storage platforms.
  • Experience operating PostgreSQL or PostGIS in Kubernetes environments.
  • Experience with distributed messaging systems.
  • Experience supporting GPU-enabled workloads in Kubernetes.
  • Familiarity with reliability engineering practices.

Benefits

  • Benefits of a dynamic and growing organization.
  • A challenging, hands-on role that will have real impact on the business.
  • Competitive compensation.
  • Comprehensive benefits package.
  • 401(k) Savings Plan.
  • Medical/Dental/Vision Benefits.
  • Health Savings Account (HSA) and Flexible Spending Account (FSA).
  • Unlimited Paid Time-off.
  • 11 Paid Holidays.
  • Paid Parental Leave.
  • Company Paid Short-term Disability (STD).
  • Company Paid Long-term Disability (LTD).
  • Company Paid Life Insurance.
Before You Apply
️
remote Be aware of the location restriction for this remote position: Australia
β€Ό Beware of scams! When applying for jobs, you should NEVER have to pay anything. Learn more.
Senior Site Reliability Engineer @Climavision
All Others
Salary $130,000 - $170..
Remote Location
Employment Type full-time
Posted 1wk ago
Apply for this position
Did not apply βœ“
Applied βœ“
Sent Follow-Up βœ“
Interview Scheduled βœ“
Interview Completed βœ“
Offer Accepted βœ“
Offer Declined βœ“
Application Denied βœ“
Unlock 125,000+ Remote Jobs
️
remote Be aware of the location restriction for this remote position: Australia
β€Ό Beware of scams! When applying for jobs, you should NEVER have to pay anything. Learn more.
Apply for this position
Did not apply βœ“
Applied βœ“
Sent Follow-Up βœ“
Interview Scheduled βœ“
Interview Completed βœ“
Offer Accepted βœ“
Offer Declined βœ“
Application Denied βœ“
Unlock 125,000+ Remote Jobs
Γ—
Apply to the best remote jobs
before everyone else

Access 125,000+ vetted remote jobs and get daily alerts.

4.9 β˜…β˜…β˜…β˜…β˜… from 500+ reviews

⚑ 126,680+ remote jobs, refreshed hourly

πŸ”” Real-time alerts: Apply first, direct to employer

πŸ›‘οΈ Vetted companies, no scams, true remote only

Unlock All Jobs Now

Maybe later