Role Description
True Zero is seeking a Datadog Observability Engineer to lead the design, implementation, and operation of enterprise-grade observability solutions built on the Datadog platform across secure, mission-critical, and healthcare environments. This is a senior technical role for an engineer who pairs deep, hands-on expertise in infrastructure monitoring, APM, log management, and synthetic testing with the judgment to guide teams, engage clients, and deliver production-grade, highly reliable observability capabilities.
The ideal candidate brings a track record of architecting and operating Datadog environments across large, distributed systems β including experience monitoring healthcare applications and infrastructure subject to HIPAA and other regulatory requirements, where uptime, data integrity, and patient-safety-critical alerting are paramount. Experience integrating endpoint and asset visibility from Tanium into the overall observability and security architecture is a strong plus, giving teams a unified view across infrastructure health, application performance, and endpoint risk.
Job Responsibilities
-
Observability Architecture:
Design, implement, and maintain end-to-end observability solutions using Datadog β infrastructure monitoring, APM/distributed tracing, log management, RUM, and synthetic monitoring β across cloud and on-premises environments.
-
Dashboard & Alerting Development:
Build and maintain dashboards, SLOs, and intelligent alerting policies that surface actionable signal and reduce noise for on-call teams.
-
Healthcare Systems Monitoring:
Implement and tune observability for healthcare applications and infrastructure (EHR/EMR platforms, HL7/FHIR interfaces, clinical systems) to protect uptime, data integrity, and compliance with HIPAA and related regulatory requirements.
-
Log & Metrics Pipeline Engineering:
Configure log pipelines, custom metrics, and telemetry ingestion (via Datadog Agent, OpenTelemetry, and APIs) to ensure high-fidelity, cost-efficient data collection at scale.
-
Incident Response & On-Call Support:
Partner with SRE and engineering teams to detect, triage, and resolve incidents; drive root-cause analysis and post-incident reviews using Datadogβs tooling.
-
Systems & Security Integration:
Integrate the Datadog platform with existing ITSM, CI/CD, and security tooling; collaborate with security teams to incorporate endpoint and asset data from Tanium into the overall observability and security architecture for unified risk and health visibility.
-
Infrastructure as Code:
Manage Datadog configuration (monitors, dashboards, SLOs) as code using Terraform or the Datadog API/CLI to ensure consistency, version control, and repeatable deployments.
-
Cost & Performance Optimization:
Monitor and optimize Datadog usage, tagging strategy, and data retention to control cost while maintaining coverage and performance.
-
Governance & Compliance:
Ensure observability practices align with regulatory and security requirements, including HIPAA, NIST SP 800-53, and Zero Trust principles.
-
Reporting:
Deliver technical reports on system health, performance trends, and reliability posture with clear, actionable recommendations.
-
Client Collaboration:
Engage with clients and stakeholders throughout the project lifecycle to align observability strategy with mission and business objectives.
-
Team Development:
Coach and mentor junior engineers on observability best practices and Datadog platform capabilities.
Qualifications
-
Minimum of 5 years in observability, DevOps, SRE, or systems engineering, with substantial hands-on experience administering and architecting the Datadog platform.
-
U.S. citizenship required; must be willing to undergo a U.S. Government background investigation.
-
Deep hands-on experience with Datadog APM, Infrastructure Monitoring, Log Management, Synthetic Monitoring, RUM, and Cloud Security Management.
-
Experience monitoring and supporting healthcare applications and infrastructure (EHR/EMR, HL7/FHIR, clinical or patient-facing systems) in HIPAA-regulated environments is highly valued.
-
Familiarity with Tanium or similar endpoint management/visibility platforms, and experience incorporating endpoint telemetry into the overall observability and security architecture, is a strong plus.
-
Hands-on experience with major cloud platforms (AWS, Azure, GCP), containerized environments (Kubernetes, Docker), and hybrid/on-prem infrastructure.
-
Proficiency with Terraform and scripting/automation (Python, Bash) for configuration-as-code and pipeline automation.
-
Strong understanding of SRE principles, SLIs/SLOs/error budgets, incident management, and root-cause analysis.
-
Familiarity with NIST SP 800-53, HIPAA technical safeguards, and Zero Trust architecture principles.
-
Experience integrating observability tooling with ITSM platforms (ServiceNow, Jira), CI/CD pipelines, and SIEM/security tooling.
-
Excellent written and verbal skills to convey technical information to diverse audiences.
-
Bachelorβs degree in Computer Science, Information Technology, Engineering, or a related field preferred; equivalent experience considered.
Certifications
-
Datadog Fundamentals Certification
-
Datadog APM & Distributed Tracing Fundamentals Certification
-
Datadog Log Management Fundamentals Certification
-
AWS Certified Solutions Architect β Associate or AWS Certified DevOps Engineer β Professional
-
Microsoft Certified: Azure Administrator Associate or Azure Solutions Architect Expert
-
Google Cloud Professional Cloud DevOps Engineer
-
Certified Kubernetes Administrator (CKA)
-
HashiCorp Certified: Terraform Associate
-
ITIL Foundation Certification
Benefits
-
Competitive salary, paid twice per month
-
Best in class medical coverage
-
100% of medical premiums covered by True Zero
-
Company wide new business incentive programs
-
Contribution Incentives (i.e. white papers, blog posts, internal webinars, etc.)
-
3 weeks of PTO starting + 11 Paid Holidays Annually
-
401k Program with 100% company match on the first 4%
-
Monthly reimbursement of Cell Phone and Home Internet costs
-
Paternity/Maternity Leave
-
Investment in training and certifications to broaden and deepen your technical skills