Data Reliability Engineer @Vytalize Health
Data and Analytics
Salary unspecified
Remote Location
πŸ‡ΊπŸ‡Έ USA Only
Employment Type full-time
Posted 2mths ago

[Hiring] Data Reliability Engineer @Vytalize Health

2mths ago - Vytalize Health is hiring a remote Data Reliability Engineer. πŸ’Έ Salary: unspecified πŸ“Location: USA

Role Description

The Data Reliability Engineer (DRE) at Vytalize Health is responsible for ensuring the end-to-end reliability, quality, and operational health of data across the full data lifecycle β€” from ingestion through downstream delivery and consumption. This role sits at the intersection of Data Engineering and Data Services, with a primary focus on building confidence that data is accurate, timely, observable, and dependable for both internal and external consumers.

The DRE role applies Site Reliability Engineering (SRE) principles to data systems, emphasizing proactive monitoring, automation, failure prevention, and rapid recovery. This individual partners closely with Data Engineering, Data Services, DevOps, Product, and Analytics teams to define and enforce reliability standards, service levels, and operational practices for mission-critical healthcare data pipelines and data products.

Given the sensitive and regulated nature of healthcare data, this role plays a key part in ensuring data reliability while maintaining strict compliance with security, privacy, and regulatory requirements. You will be metrics-driven β€” establishing clear reliability targets (SLIs, SLOs, SLAs) and measuring success through data quality, freshness, and delivery timeliness.

Qualifications

  • Bachelor's degree in Computer Science, Engineering, Information Systems, or a related field, or equivalent professional experience.
  • 5+ years of experience working with data platforms, data pipelines, or distributed data systems in production environments.
  • Demonstrated experience improving reliability, observability, or operational quality of data systems with measurable SLI/SLO/SLA improvements.
  • Hands-on experience supporting both data ingestion pipelines and downstream data consumption or delivery patterns.
  • 1+ years of hands-on experience with machine learning-based monitoring, anomaly detection, or AI-assisted observability tools.
  • Demonstrated experience with data quality testing, validation frameworks, and quality metrics definition.

Requirements

  • Strong understanding of modern data architectures, including data lakehouse patterns and multi-layer (bronze/silver/gold) data models.
  • Experience with cloud-based data platforms (AWS, Databricks, or similar).
  • Proficiency in Python and SQL, with experience building or supporting production-grade data pipelines.
  • Experience implementing data quality frameworks, monitoring tools, and alerting systems.
  • Demonstrated expertise with workflow orchestration tools (e.g., Databricks Workflows, Airflow) and version-controlled deployment practices.
  • Familiarity with SRE and reliability engineering concepts including SLIs, SLOs, error budgets, and blameless postmortem culture.
  • Strong troubleshooting and root cause analysis skills across complex, distributed systems.
  • Experience designing and operating observability systems for data pipelines (metrics, logs, traces, alerts).
  • Ability to communicate clearly with both technical and non-technical stakeholders during incidents, postmortems, and requirements discussions.
  • Understanding of healthcare data, EMR integrations, or regulated data environments is strongly preferred.
  • Experience defining and measuring data quality metrics; ability to establish and track reliability KPIs.

Benefits

  • Hands-on experience with ML-based anomaly detection frameworks or tools (e.g., Datadog Anomaly Detection, cloud-native monitoring ML, custom model development).
  • Experience leveraging LLMs or AI-assisted tools (e.g., Claude Code, ChatGPT, GitHub Copilot) to accelerate development of monitoring code, incident response workflows, and documentation.
  • Familiarity with healthcare data standards: FHIR, HL7, CCD, claims data formats, and value-based care metrics.
  • Experience operating observability and incident management platforms (e.g., DataDog, New Relic, Sumo Logic, PagerDuty).
  • On-call experience and demonstrated comfort with incident response, runbook creation, and blameless postmortem analysis.
  • Experience with policy-as-code and data governance frameworks.
  • Background in a startup or high-growth environment with exposure to scaling data systems.
  • Familiarity with Tuva or similar clinical data normalization and quality frameworks.
Before You Apply
️
πŸ‡ΊπŸ‡Έ Be aware of the location restriction for this remote position: USA Only
β€Ό Beware of scams! When applying for jobs, you should NEVER have to pay anything. Learn more.
Data Reliability Engineer @Vytalize Health
Data and Analytics
Salary unspecified
Remote Location
πŸ‡ΊπŸ‡Έ USA Only
Employment Type full-time
Posted 2mths ago
Apply for this position
Did not apply βœ“
Applied βœ“
Sent Follow-Up βœ“
Interview Scheduled βœ“
Interview Completed βœ“
Offer Accepted βœ“
Offer Declined βœ“
Application Denied βœ“
Unlock 125,000+ Remote Jobs
️
πŸ‡ΊπŸ‡Έ Be aware of the location restriction for this remote position: USA Only
β€Ό Beware of scams! When applying for jobs, you should NEVER have to pay anything. Learn more.
Apply for this position
Did not apply βœ“
Applied βœ“
Sent Follow-Up βœ“
Interview Scheduled βœ“
Interview Completed βœ“
Offer Accepted βœ“
Offer Declined βœ“
Application Denied βœ“
Unlock 125,000+ Remote Jobs
Γ—
Apply to the best remote jobs
before everyone else

Access 125,000+ vetted remote jobs and get daily alerts.

4.9 β˜…β˜…β˜…β˜…β˜… from 500+ reviews

⚑ 127,048+ remote jobs, refreshed hourly

πŸ”” Real-time alerts: Apply first, direct to employer

πŸ›‘οΈ Vetted companies, no scams, true remote only

Unlock All Jobs Now

Maybe later