Healthcare Data Scientist @Cherokee Federal
Data and Analytics
Salary unspecified
Remote Location
πŸ‡ΊπŸ‡Έ USA Only
Employment Type full-time
Posted 1wk ago

[Hiring] Healthcare Data Scientist @Cherokee Federal

1wk ago - Cherokee Federal is hiring a remote Healthcare Data Scientist. πŸ’Έ Salary: unspecified πŸ“Location: USA

Role Description

ATA is seeking a Data Scientist to support data pipeline development, validation, and analysis within a cloud-based Health IT data platform. This role is hands-on and delivery-focused, with an emphasis on building reliable, reproducible data workflows using SQL, Python, and PySpark in an Azure Synapse environment.

A core expectation of this role is the ability to work across the full data lifecycle, from ingestion through transformation to final dataset delivery, while maintaining data quality and traceability. The ideal candidate is comfortable debugging data issues end-to-end, understands how data structure and join logic impact outputs, and applies disciplined validation and documentation practices. This role also supports exploratory data analysis and the development of derived datasets to enable analytics and downstream use cases. The position will work extensively with healthcare data originating from EHR systems and interface feeds, including HL7 v2 and FHIR data, clinical terminology, and source-to-target data mappings.

Qualifications

  • Hands-on experience writing SQL queries for data transformation and analysis.
  • Experience using Python (e.g., pandas) for data processing.
  • Hands-on experience using PySpark for distributed data processing.
  • Experience working within cloud-based data platforms, preferably Azure Synapse or similar.
  • Understanding ETL/ELT concepts and data pipeline architecture.
  • Experience working with structured and semi-structured data formats (CSV, JSON, Parquet).
  • Familiarity with Git and collaborative development workflows.
  • Strong problem-solving and debugging skills across data pipelines.
  • Ability to validate and ensure data quality through structured checks and testing practices.
  • Strong written and verbal communication skills.

Requirements

  • Develop, maintain, and optimize data pipelines using SQL, Python (pandas), and PySpark.
  • Execute and manage notebook-based workflows within Azure Synapse, including debugging and documentation.
  • Process and transform structured and semi-structured data in formats such as CSV, JSON/NDJSON, and Parquet.
  • Work within ETL/ELT pipelines across raw, curated, and production data layers.
  • Ingest, profile, map, and transform healthcare data from EHR systems and interface feeds while preserving source lineage and clinical context.
  • Perform structured data validation, including row counts, null checks, duplicate detection, schema validation, and allowed value enforcement.
  • Identify and resolve data quality issues such as schema drift, inconsistencies, and transformation errors across pipeline stages.
  • Apply repeatable testing and validation practices, including reproducing issues, verifying fixes, and ensuring data reliability prior to downstream use.
  • Validate source-to-target mappings and reconcile records across source and destination systems during data conversion and migration activities.
  • Conduct exploratory data analysis to identify patterns, anomalies, and data quality concerns.
  • Develop derived datasets to support reporting, analytics, and downstream data use cases.
  • Collaborate with stakeholders to translate data requirements into usable datasets and metrics.
  • Interpret and work with data schemas, including column definitions, data types, primary and composite keys, and table relationships.
  • Manage dataset grain and understand how join strategies (e.g., one-to-one vs. one-to-many) impact row counts and outputs.
  • Trace data issues from source ingestion through transformation logic to final outputs.
  • Use logs and debugging approaches to diagnose and resolve pipeline issues.
  • Document data transformations, assumptions, mappings, and validation results in a clear and consistent manner.
  • Collaborate with engineers, analysts, and stakeholders to ensure data usability, integrity, and alignment with requirements.
  • Communicate data issues, findings, and workflow updates with technical team members.

Benefits

  • Generous paid time-off.
  • Employee incentive program.
  • Continuous learning culture, Internal Investment Projects (IIP), virtual brown-bags/level-ups, and other professional development activities.
  • Recruiting bonuses.
  • 3% 401k Safe Harbor contributions.
  • Medical/Dental/Vision, Long & Short-term Disability, AD&D insurance, and Life Insurance.
Before You Apply
️
πŸ‡ΊπŸ‡Έ Be aware of the location restriction for this remote position: USA Only
β€Ό Beware of scams! When applying for jobs, you should NEVER have to pay anything. Learn more.
Healthcare Data Scientist @Cherokee Federal
Data and Analytics
Salary unspecified
Remote Location
πŸ‡ΊπŸ‡Έ USA Only
Employment Type full-time
Posted 1wk ago
Apply for this position
Did not apply βœ“
Applied βœ“
Sent Follow-Up βœ“
Interview Scheduled βœ“
Interview Completed βœ“
Offer Accepted βœ“
Offer Declined βœ“
Application Denied βœ“
Unlock 120,000+ Remote Jobs
️
πŸ‡ΊπŸ‡Έ Be aware of the location restriction for this remote position: USA Only
β€Ό Beware of scams! When applying for jobs, you should NEVER have to pay anything. Learn more.
Apply for this position
Did not apply βœ“
Applied βœ“
Sent Follow-Up βœ“
Interview Scheduled βœ“
Interview Completed βœ“
Offer Accepted βœ“
Offer Declined βœ“
Application Denied βœ“
Unlock 120,000+ Remote Jobs
Γ—

Apply to the best remote jobs
before everyone else

Access 120,000+ vetted remote jobs and get daily alerts.

4.9 β˜…β˜…β˜…β˜…β˜… from 500+ reviews
Unlock All Jobs Now

Maybe later