Director of Site Reliability Engineering @ServiceNow
All Others
Salary $221,200 - $387..
Remote Location
πŸ‡ΊπŸ‡Έ USA Only
Employment Type full-time
Posted 2d ago

[Hiring] Director of Site Reliability Engineering @ServiceNow

2d ago - ServiceNow is hiring a remote Director of Site Reliability Engineering. πŸ’Έ Salary: $221,200 - $387,100 per year πŸ“Location: USA

Role Description

We are looking for a Director of Site Reliability Engineering to lead the next phase of our reliability transformation as ServiceNow modernizes toward a cloud-agnostic, cloud-ready production platform. This leader will own key elements of the SRE operating model across:

  • Reliability Engineering
  • Service Enablement
  • Service Registry
  • SLI/SLO standards
  • Reliability governance
  • Automation
  • AI-enabled operations
  • Production readiness

The role will lead a global engineering organization and partner across Product Engineering, Infrastructure, Architecture, Security, Release Engineering, and Customer Support to establish consistent reliability practices across ServiceNow products and services. The Director will play a critical role in evolving the organization from reactive operations toward an engineering-led SRE model focused on prevention, automation, resilience, and continuous improvement.

What you get to do in this role:

  • Define and execute the SRE strategy and operating model across reliability engineering, service enablement, observability, automation, incident learning, and production readiness.
  • Lead and develop a global organization of engineering managers, technical leaders, and SREs.
  • Establish enterprise reliability standards for service ownership, tiering, golden signals, SLIs/SLOs, error budgets, alerting, on-call practices, and service health reviews.
  • Lead the Service Enablement strategy by establishing minimum reliability requirements and maturity standards for critical services.
  • Own the Service Registry strategy, improving service ownership, dependency visibility, maturity tracking, and impact-aware operational decision-making.
  • Drive adoption of SLIs, SLOs, error budgets, and burn-rate alerting across critical services, ensuring teams consistently use reliability signals to manage customer impact.
  • Build a culture of engineering away toil by turning recurring operational work and incident patterns into automation, self-service, and systemic fixes.
  • Establish the AI-enabled SRE roadmap, including change-risk assessment, operational insights, remediation recommendations, and policy-driven automation.
  • Drive reliability and production-readiness strategy across AWS, Azure, and GCP by establishing cloud-agnostic patterns while addressing hyperscaler-specific operational requirements.
  • Partner with product and platform engineers to design, launch, and operate reliable services throughout the production lifecycle.
  • Establish launch and production-readiness practices that validate availability, latency, performance, capacity, dependencies, rollback, and recovery before customer impact.
  • Drive sustainable operations by scaling self-service capabilities, automation platforms, and systemic reliability improvements across engineering teams.
  • Lead incident response, blameless postmortems, and corrective actions that convert production failures into lasting reliability improvements.
  • Measure reliability through SLIs, SLOs, error budgets, golden signals, change failure rate, MTTR, capacity health, and toil reduction.
  • Influence architecture and platform direction to simplify operating models and improve reliability across ServiceNow's global infrastructure.
  • Partner with executive and engineering leaders to prioritize reliability investments and drive adoption beyond the direct SRE organization.

Qualifications

  • Experience in leveraging or critically thinking about how to integrate AI into work processes, decision-making, or problem-solving.
  • 12 years of significant leadership experience in Site Reliability Engineering, Production Engineering, Platform Engineering, Cloud Infrastructure, or large-scale distributed systems with a Bachelor's degree; or 8 years and a Master's degree; or a PhD with 5 years experience; or equivalent experience.
  • Proven success leading managers and senior technical leaders across geographically distributed engineering organizations.
  • Demonstrated success leading SRE, infrastructure, or reliability transformation at scale.
  • Strong understanding of SLIs/SLOs, error budgets, observability, incident management, reliability governance, and on-call practices.
  • Experience with service catalogs, service registries, service ownership models, Backstage, CMDB, dependency mapping, or service topology.
  • Strong background in cloud infrastructure and modernization across AWS, Azure, and/or GCP.
  • Understanding of Kubernetes, distributed systems, networking, databases, infrastructure automation, and cloud-native architecture.
  • Experience driving automation through orchestration, Infrastructure as Code, self-service platforms, and auto-remediation.
  • Familiarity with AI-assisted operations, autonomous remediation, or agentic technologies is highly desirable.
  • Experience establishing production-readiness practices for releases, resilience, disaster recovery, infrastructure changes, and cloud migrations.
  • Ability to use incident, reliability, and operational data to prioritize engineering work and drive systemic improvements.
  • Strong cross-functional influence and executive communication skills.
  • Ability to operate effectively through ambiguity, organizational transformation, and large-scale technical change.

Requirements

  • For positions in this location, we offer a base pay of $221,200 - $387,100, plus equity (when applicable), variable/incentive compensation and benefits.
  • Sales positions generally offer a competitive On Target Earnings (OTE) incentive compensation structure.
  • Please note that the base pay shown is a guideline, and individual total compensation will vary based on factors such as qualifications, skill level, competencies, and work location.
  • We also offer health plans, including flexible spending accounts, a 401(k) Plan with company match, ESPP, matching donations, a flexible time away plan and family leave programs.

Benefits

  • Flexible spending accounts
  • 401(k) Plan with company match
  • Employee Stock Purchase Plan (ESPP)
  • Matching donations
  • Flexible time away plan
  • Family leave programs
Before You Apply
️
πŸ‡ΊπŸ‡Έ Be aware of the location restriction for this remote position: USA Only
β€Ό Beware of scams! When applying for jobs, you should NEVER have to pay anything. Learn more.
Director of Site Reliability Engineering @ServiceNow
All Others
Salary $221,200 - $387..
Remote Location
πŸ‡ΊπŸ‡Έ USA Only
Employment Type full-time
Posted 2d ago
Apply for this position
Did not apply βœ“
Applied βœ“
Sent Follow-Up βœ“
Interview Scheduled βœ“
Interview Completed βœ“
Offer Accepted βœ“
Offer Declined βœ“
Application Denied βœ“
Unlock 125,000+ Remote Jobs
️
πŸ‡ΊπŸ‡Έ Be aware of the location restriction for this remote position: USA Only
β€Ό Beware of scams! When applying for jobs, you should NEVER have to pay anything. Learn more.
Apply for this position
Did not apply βœ“
Applied βœ“
Sent Follow-Up βœ“
Interview Scheduled βœ“
Interview Completed βœ“
Offer Accepted βœ“
Offer Declined βœ“
Application Denied βœ“
Unlock 125,000+ Remote Jobs
Γ—

Apply to the best remote jobs
before everyone else

Access 125,000+ vetted remote jobs and get daily alerts.

4.9 β˜…β˜…β˜…β˜…β˜… from 500+ reviews

⚑ 128,326+ remote jobs, refreshed hourly

πŸ”” Real-time alerts: Apply first

πŸ›‘οΈ Vetted companies, no scams, true remote only

Unlock All Jobs Now

Maybe later