Role Description
We are seeking an experienced Senior Site-Reliability Engineer to join our infrastructure team. The ideal candidate will be responsible for ensuring the reliability, availability, and performance of our Windows-based production environments. You will bridge development and operations to deliver highly available services while maintaining operational excellence.
-
Design, implement, and maintain scalable infrastructure using Infrastructure as Code (IaC) practices
-
Develop and maintain automation scripts using various scripting languages (PowerShell, Python, Ruby, etc) for OS provisioning, configuration management, and operational tasks
-
Implement and manage configuration management solutions (Terraform, Puppet, and/or Chef) across hybrid environments
-
Monitor system health, performance, and availability using industry-standard tools and practices
-
Establish and enforce SLAs, SLOs, and error budgets for production services
-
Participate in on-call rotation and respond to incidents with a focus on rapid restoration and root cause analysis
-
Collaborate with development teams to improve deployment pipelines and release processes
-
Document operational procedures, runbooks, and architectural decisions
-
Conduct post-mortem reviews and implement corrective actions to prevent recurrence
Qualifications
-
5+ years of experience in Systems Administration, DevOps, or Site-Reliability Engineering roles
-
Strong expertise in Windows Server environments (2016+), including Active Directory, IIS, and MS SQL
-
Strong cloud skills, AWS experience preferred
-
Advanced proficiency in scripting, including module development and integration with REST APIs
-
Hands-on experience with configuration management tools:
-
Terraform (infrastructure provisioning)
-
Puppet or Chef (configuration management)
-
Experience with monitoring and observability platforms (e.g., Prometheus, Grafana, Datadog, New Relic)
-
Solid understanding of networking concepts (DNS, TCP/IP, Load Balancing, VPN)
-
Strong problem-solving skills with the ability to troubleshoot complex issues across multiple technology layers
Requirements
-
Bachelor's degree in Computer Science, Information Technology, or related field (or equivalent professional experience)
-
Certifications such as AWS Solutions Architect, Microsoft Certifications, or HashiCorp Certified: Terraform Associate
-
Experience with containerization technologies (Docker, Kubernetes)
-
Familiarity with CI/CD tools (Gitlab Pipelines, Jenkins, GitHub Actions)
-
Knowledge of security best practices and compliance frameworks
-
Experience with log aggregation and analysis tools (ELK Stack, Splunk)
Benefits
-
Comprehensive benefits package including medical, dental, and vision insurance
-
Life, AD&D, and disability insurance
-
Paid time off and 11 company holidays
-
401(k) retirement plan with company matching
-
Additional employee benefits and wellness resources