Enterprise Service Reliability and Insights Lead @DecisionPoint | Cortek
All Others
Salary unspecified
Remote Location
πŸ‡ΊπŸ‡Έ USA Only
Employment Type full-time
Posted 2mths ago

[Hiring] Enterprise Service Reliability and Insights Lead @DecisionPoint | Cortek

2mths ago - DecisionPoint | Cortek is hiring a remote Enterprise Service Reliability and Insights Lead. πŸ’Έ Salary: unspecified πŸ“Location: USA

Role Description

DecisionPoint seeks a Senior Enterprise Service Reliability and Insights Lead to oversee enterprise-wide monitoring, observability, and operational intelligence for a large federal and DoD-aligned IT environment. This senior-level role defines the monitoring strategy, manages toolsets, develops dashboards, establishes alerting thresholds, and ensures service reliability through proactive detection and rapid incident identification.

The Enterprise Service Reliability and Insights Lead is responsible for driving visibility into uptime, system performance, service health, and operational risks. This position partners closely with Tier 2 and Tier 3 engineering teams, cloud operations, cybersecurity, and service desk leadership to ensure monitoring aligns with mission needs, SLAs, and enterprise performance objectives. This position is fully remote.

Duties & Responsibilities

  • Define, implement, and manage the enterprise monitoring and observability strategy.
  • Oversee monitoring tools, dashboards, agents, log pipelines, and alerting configurations across all environments.
  • Establish alert thresholds, escalation criteria, and performance indicators that support proactive issue detection.
  • Ensure monitoring coverage aligns with uptime, performance, and security requirements.
  • Collaborate with Tier 2 and Tier 3 engineering teams on system health assessments, log analytics, and incident triage.
  • Lead efforts to correlate events across application, infrastructure, network, and security monitoring tools.
  • Deliver actionable insights on system reliability, capacity issues, performance bottlenecks, and incident trends.
  • Support SLA and KPI measurement, reporting, and compliance tracking.
  • Maintain monitoring documentation, dashboards, service health definitions, and alerting standards.
  • Partner with cloud, infrastructure, and cybersecurity teams to ensure observability supports mission and compliance needs.
  • Recommend improvements to monitoring architectures, event correlation, and automation capabilities.
  • Participate in incident response activities, root cause analysis sessions, and readiness reviews.
  • Drive continuous improvement initiatives across reliability engineering and service monitoring.

Qualifications

  • Must hold an active Secret clearance, supported by a Tier 3 background investigation.
  • Bachelor’s degree in Information Technology, Cybersecurity, Systems Engineering, or a related technical field.
  • Minimum 10 years of experience in service reliability, monitoring engineering, IT operations, or systems engineering.
  • Experience designing or managing enterprise monitoring systems and dashboards.
  • Experience defining SLAs, KPIs, and operational performance measurements.
  • Experience collaborating with Tier 2 and Tier 3 teams for incident management and problem resolution.
  • Experience with log analysis, event correlation, and observability platforms.
  • Strong understanding of monitoring and observability tools (metrics, logs, traces).
  • Knowledge of uptime, performance, and reliability engineering practices.
  • Familiarity with ITIL v4 processes for incident, problem, and change management.
  • Understanding of alerting strategies, threshold design, and escalation workflows.
  • Knowledge of DoD or federal IT operational environments.
  • Experience with cloud-native monitoring services and distributed systems monitoring.
  • Experience with APM tools, SIEM integrations, or event correlation engines.
  • Familiarity with automation scripting or analytics for monitoring enhancement.
  • Required certifications: ITIL v4 Foundation, CompTIA Security+.
  • Preferred certifications: Cloud monitoring certifications (AWS, Azure, or similar), SRE or observability-related certifications.

Skills

  • Strong analytical skills for interpreting system health and service reliability data.
  • Excellent communication and reporting skills for executive and technical audiences.
  • Ability to lead cross-functional coordination during performance events and incidents.
  • High attention to detail with strong documentation habits.
  • Ability to drive continuous improvement across monitoring, reliability, and availability functions.
Before You Apply
️
πŸ‡ΊπŸ‡Έ Be aware of the location restriction for this remote position: USA Only
β€Ό Beware of scams! When applying for jobs, you should NEVER have to pay anything. Learn more.
Enterprise Service Reliability and Insights Lead @DecisionPoint | Cortek
All Others
Salary unspecified
Remote Location
πŸ‡ΊπŸ‡Έ USA Only
Employment Type full-time
Posted 2mths ago
Apply for this position
Did not apply βœ“
Applied βœ“
Sent Follow-Up βœ“
Interview Scheduled βœ“
Interview Completed βœ“
Offer Accepted βœ“
Offer Declined βœ“
Application Denied βœ“
Unlock 125,000+ Remote Jobs
️
πŸ‡ΊπŸ‡Έ Be aware of the location restriction for this remote position: USA Only
β€Ό Beware of scams! When applying for jobs, you should NEVER have to pay anything. Learn more.
Apply for this position
Did not apply βœ“
Applied βœ“
Sent Follow-Up βœ“
Interview Scheduled βœ“
Interview Completed βœ“
Offer Accepted βœ“
Offer Declined βœ“
Application Denied βœ“
Unlock 125,000+ Remote Jobs
Γ—
Apply to the best remote jobs
before everyone else

Access 125,000+ vetted remote jobs and get daily alerts.

4.9 β˜…β˜…β˜…β˜…β˜… from 500+ reviews

⚑ 126,868+ remote jobs, refreshed hourly

πŸ”” Real-time alerts: Apply first, direct to employer

πŸ›‘οΈ Vetted companies, no scams, true remote only

Unlock All Jobs Now

Maybe later