Director, Data & Storage Reliability Engineering @ServiceNow
All Others
Salary usd 221,200 - 3..
Remote Location
πŸ‡ΊπŸ‡Έ USA Only
Employment Type full-time
Posted 3wks ago

[Hiring] Director, Data & Storage Reliability Engineering @ServiceNow

3wks ago - ServiceNow is hiring a remote Director, Data & Storage Reliability Engineering. πŸ’Έ Salary: usd 221,200 - 387,100 per year πŸ“Location: USA

Role Description

The successful candidate will lead the Data & Storage Reliability Engineering organization responsible for improving reliability, resilience, performance, scalability, observability, and customer experience across ServiceNow's database, storage, and supporting platform infrastructure.

  • Build, develop, and scale high-performing engineering teams focused on:
    • Reliability engineering
    • Observability
    • Performance engineering
    • Diagnostics
    • Automation
    • Production analytics
    • Migration readiness
    • Resilience engineering
    • Prevention engineering
  • Responsibilities include:
    • Talent acquisition
    • Performance management
    • Career development
    • Succession planning
    • Objective setting
    • Coaching
    • Prioritization of strategic initiatives
  • Establish a strong engineering-first culture centered on:
    • Data-driven decision making
    • Continuous improvement
    • Operational excellence
    • Customer experience
    • Systemic risk reduction
  • Accountable for identifying:
    • Recurring failure patterns
    • Reliability risks
    • Performance bottlenecks
    • Scalability constraints
    • Migration challenges
    • Operational inefficiencies
  • Partner closely with:
    • Product Engineering
    • Database Engineering
    • Cloud Infrastructure
    • Architecture
    • Storage Engineering
    • Support
    • Operations teams
  • Serve as the senior technical leader for:
    • Complex reliability investigations
    • Customer-critical escalation reviews
    • Migration readiness assessments
    • Platform improvement initiatives
  • Influence architectural decisions and technology investments by providing:
    • Reliability expertise
    • Observability insights
    • Performance guidance
    • Production-based evidence
  • Establish scalable reliability engineering practices, standards, governance processes, and operating models across the organization.
  • Drive adoption of:
    • Observability standards
    • Reliability engineering frameworks
    • Resiliency assessments
    • Migration readiness practices
    • Diagnostics capabilities
    • Engineering guardrails
    • Automation strategies
  • Continuously evaluate incidents, customer escalations, migration outcomes, platform telemetry, performance trends, capacity signals, and operational data to identify systemic risks and drive long-term engineering improvements.
  • Establish a formal review process with SWAT and Customer & Production Engineering teams to evaluate major incidents, recurring operational challenges, migration learnings, customer-impacting events, and emerging platform risks.
  • Establish meaningful KPIs and engineering metrics that provide visibility into:
    • Platform reliability
    • Resiliency
    • Performance
    • Operational efficiency
    • Customer experience
    • Engineering productivity
    • Risk reduction
  • Leverage AI-powered tools, analytics, automation frameworks, and production intelligence to identify emerging risks, improve detection coverage, accelerate engineering insights, reduce operational toil, and improve engineering productivity.
  • Maintain a portfolio of reliability investments spanning:
    • Observability
    • Performance
    • Diagnostics
    • Resilience
    • Automation
    • Prevention
  • Champion a proactive reliability engineering model that shifts the organization from reactive issue response toward predictive analysis, prevention, resilience, and continuous optimization.

Qualifications

  • Experience in leveraging or critically thinking about how to integrate AI into work processes, decision-making, or problem-solving.
  • Strong product mindset with demonstrated experience treating technical capabilities as products.
  • Experience translating production insights, customer pain points, operational challenges, reliability risks, and platform telemetry into prioritized engineering investments.
  • Experience partnering closely with production operations, customer escalation teams, reliability organizations, and software engineering teams.
  • Experience defining product strategies, developing roadmaps, prioritizing investments, and aligning stakeholders across multiple organizations.
  • Experience operating a portfolio of engineering investments.
  • 15+ years of experience in software engineering, platform engineering, reliability engineering, infrastructure engineering, database engineering, distributed systems, product management, or large-scale SaaS environments.
  • 8+ years of engineering leadership experience, including leading managers and globally distributed teams.
  • Extensive experience leading Reliability Engineering, Platform Engineering, Database Engineering, Infrastructure Engineering, Production Engineering, Performance Engineering, or related technical organizations.
  • Deep expertise in distributed systems, databases, storage technologies, cloud infrastructure, and large-scale SaaS architectures.
  • Strong understanding of reliability engineering principles, observability, scalability, resiliency, operational excellence, and performance engineering.
  • Experience building and operating observability, telemetry, diagnostics, reliability, or performance capabilities at scale.
  • Proven experience identifying systemic issues and converting operational insights into strategic engineering improvements.
  • Experience partnering closely with Product Management organizations.
  • Exceptional communication, stakeholder management, and leadership skills.
  • Bachelor's degree in Computer Science, Engineering, or a related technical field, or equivalent practical experience.

Requirements

  • Previous Product Management experience in a platform, infrastructure, cloud, database, storage, or SaaS environment.
  • Experience applying product management disciplines such as roadmap planning, prioritization, customer-centric thinking, outcome measurement, and portfolio management to engineering organizations.
  • Experience operating large-scale enterprise database and storage platforms supporting mission-critical workloads.
  • Experience building and scaling Reliability Engineering, Performance Engineering, Platform Engineering, SRE, or Production Engineering organizations.
  • Experience with observability platforms, telemetry systems, diagnostics frameworks, and production analytics.
  • Experience with migration readiness, resiliency validation, reliability testing, operational risk reduction, and large-scale cloud transformations.
  • Experience leveraging AI technologies to improve anomaly detection, forecasting, incident analysis, prioritization, and engineering productivity.
  • Strong understanding of distributed systems architecture, cloud platform operations, and hyperscale environments.
  • Experience developing executive-facing reliability scorecards, engineering metrics, and business impact reporting.
  • Experience influencing platform architecture, database strategy, storage strategy, and long-term engineering roadmaps.
  • Experience with Linux-based production environments and large-scale cloud infrastructure.
  • Experience supporting enterprise database technologies such as MySQL, MariaDB, PostgreSQL, Oracle, SQL Server, or cloud-native database platforms.
  • Familiarity with ServiceNow platform architecture and large-scale SaaS operations.

Benefits

  • Base pay of $221,200 - $387,100, plus equity (when applicable), variable/incentive compensation, and benefits.
  • Health plans, including flexible spending accounts.
  • 401(k) Plan with company match.
  • Employee Stock Purchase Plan (ESPP).
  • Matching donations.
  • Flexible time away plan and family leave programs.
Before You Apply
️
πŸ‡ΊπŸ‡Έ Be aware of the location restriction for this remote position: USA Only
β€Ό Beware of scams! When applying for jobs, you should NEVER have to pay anything. Learn more.
Director, Data & Storage Reliability Engineering @ServiceNow
All Others
Salary usd 221,200 - 3..
Remote Location
πŸ‡ΊπŸ‡Έ USA Only
Employment Type full-time
Posted 3wks ago
Apply for this position
Did not apply βœ“
Applied βœ“
Sent Follow-Up βœ“
Interview Scheduled βœ“
Interview Completed βœ“
Offer Accepted βœ“
Offer Declined βœ“
Application Denied βœ“
Unlock 125,000+ Remote Jobs
️
πŸ‡ΊπŸ‡Έ Be aware of the location restriction for this remote position: USA Only
β€Ό Beware of scams! When applying for jobs, you should NEVER have to pay anything. Learn more.
Apply for this position
Did not apply βœ“
Applied βœ“
Sent Follow-Up βœ“
Interview Scheduled βœ“
Interview Completed βœ“
Offer Accepted βœ“
Offer Declined βœ“
Application Denied βœ“
Unlock 125,000+ Remote Jobs
Γ—
Apply to the best remote jobs
before everyone else

Access 125,000+ vetted remote jobs and get daily alerts.

4.9 β˜…β˜…β˜…β˜…β˜… from 500+ reviews

⚑ 127,028+ remote jobs, refreshed hourly

πŸ”” Real-time alerts: Apply first, direct to employer

πŸ›‘οΈ Vetted companies, no scams, true remote only

Unlock All Jobs Now

Maybe later