Role Description
Everforth ECS is seeking a talented Senior Site Reliability Engineer (SRE) to play a key role in defining, implementing, and growing our SRE practice to ensure the reliability, availability, and performance of our critical production environments.
-
Contribute to a culture of continuous improvement, identifying areas for enhancement, and driving initiatives to improve system reliability, scalability, and efficiency.
-
Demonstrated hands-on experience designing, implementing, and maintaining solutions to ensure that systems, including infrastructure and applications, are resilient, highly available, and performant.
-
Define and measure the Service Level Objectives (SLOs) and Service Level Indicators (SLIs) for our solution.
-
Set up comprehensive logging, monitoring, and alerting solutions using the Elastic stack and other tools as necessary to ensure the continuous performance of services.
-
Respond to incidents, perform root cause analyses, and implement solutions to prevent recurrences.
-
Work in close collaboration with other SRE team members, developers, testers, infrastructure engineers, DevOps engineers, and other stakeholders to integrate reliability and observability into the software development lifecycle.
Qualifications
-
Must be a US citizen with the ability to obtain Public Trust Suitability.
-
6+ years of experience as a Site Reliability Engineer (SRE) or equivalent.
-
6+ years of demonstrated experience designing, implementing, and maintaining observability solutions to include logging, monitoring, and alerting.
-
6+ years of hands-on experience with SRE tools (e.g., Elastic, Prometheus, Grafana, Splunk, etc.).
-
3+ years defining and measuring SLOs and SLIs.
-
3+ years of relevant experience using cloud platforms (AWS GovCloud preferred).
-
3+ years of hands-on programming or scripting (e.g., Python, Bash, etc.).
-
Strong knowledge of microservices, containerization, and orchestration tools (Docker, Kubernetes).
-
Proven ability to collaborate with cross-functional teams (development, testing, and product) to integrate reliability and observability into the software development lifecycle.
-
Strong problem-solving and analytical skills.
-
Proactive, detail-oriented approach to identifying inefficiencies and implementing improvements.
-
Proficient in developing Synthetic monitoring scripts using TypeScript.
Requirements
-
Salary Range: $118,000 - $177,000
Benefits
-
General Description of Benefits