Role Description
The Lead Systems Engineer is a senior individual contributor responsible for the reliability, scalability, and modernization of Intellum's platform infrastructure. This is a highly hands-on role with significant ownership across infrastructure architecture, cloud environments, deployment systems, observability, and platform reliability.
-
Own and drive key infrastructure modernization initiatives, including the continued evolution from legacy compute environments toward modern, container-orchestrated infrastructure while maintaining reliable service for enterprise customers.
-
Design and maintain infrastructure as code across multiple cloud providers, ensuring infrastructure decisions support portability, maintainability, and long-term scalability.
-
Improve the reliability and maturity of Intellum's CI/CD systems and deployment tooling so releases are efficient, observable, and recoverable.
-
Provide technical leadership across the Systems Engineering team through mentorship, architecture guidance, knowledge sharing, and support for strong engineering practices.
-
Establish and evolve SLI and SLO practices, along with the monitoring, alerting, and load-testing capabilities needed to support platform reliability.
-
Participate in and provide leadership during platform incidents, including troubleshooting, root cause analysis, and follow-through on corrective actions.
-
Drive visibility into cloud infrastructure costs and incorporate cost considerations into architecture and infrastructure decisions.
-
Improve developer experience by evolving the infrastructure and tooling engineers depend on, including development environments, deployment workflows, and production feedback loops.
-
Partner closely with Security and Engineering teams on access controls, infrastructure hardening, compliance requirements, and secure infrastructure practices.
-
Identify operational and infrastructure risks early, recommend priorities, and help drive the technical roadmap for the Systems Engineering function.
-
Contribute to the continued development of the Systems Engineering team and function, including mentoring engineers and helping build strong technical practices as the organization evolves.
-
Perform other duties as assigned.
Qualifications
-
8+ years of hands-on experience in infrastructure, DevOps, platform engineering, site reliability engineering, or a related discipline, including experience building and operating production systems.
-
Deep hands-on experience designing, operating, and troubleshooting highly available production infrastructure.
-
Production experience across more than one major cloud provider, with depth in at least one of AWS or Google Cloud and working fluency in the other.
-
Significant experience with container orchestration and Kubernetes in production environments, including cluster operations, workload configuration, reliability, and troubleshooting.
-
Experience modernizing production infrastructure, including migrations from VM-based or legacy environments toward containerized or cloud-native architectures.
-
Strong infrastructure-as-code experience using Terraform or comparable tooling, with an emphasis on repeatability and automation.
-
Experience building, operating, or significantly improving CI/CD systems and deployment infrastructure.
-
Strong incident response and troubleshooting capabilities, including experience diagnosing complex distributed-system failures and contributing to effective post-incident review.
-
Strong Linux administration skills and scripting or programming ability in Ruby, Python, or a comparable language.
-
Experience working in a SaaS environment where reliability, availability, and production stability are critical.
-
Ability to collaborate effectively with distributed teams across US and European time zones and participate in an on-call rotation.
-
Strong communication skills and the ability to provide technical direction, mentor other engineers, and influence infrastructure decisions across teams.
Requirements
-
Experience operating production infrastructure across both AWS and Google Cloud simultaneously.
-
Prior experience leading or managing engineers, whether through formal people management, technical leadership, or mentorship.
-
Experience developing engineers and helping build strong, high-performing technical teams.
-
Experience with cloud cost management or FinOps practices at meaningful scale.
-
Experience managing deployment platforms such as Spinnaker, Jenkins, or comparable tooling.
-
Experience operating a Ruby on Rails enterprise application or comparable production codebase.
-
Familiarity with SOC 2 or similar compliance frameworks and customer-facing security requirements.
-
Working knowledge of AI-assisted development tooling and its infrastructure implications.
-
Prior people leadership or management experience in a player-coach capacity, balancing hands-on technical contribution with mentorship, team guidance, and development of engineers.
-
Background in learning management systems, learning technologies, or adult education platforms.
Benefits
-
Medical - 100% of employee premiums for selected individual plans
-
Dental - 100% of employee premiums covered
-
Vision - 100% of employee premiums covered
-
LinkedIn Learning
-
401(k) plus matching (US Based Only)
-
Flexible PTO
-
Calm subscription
-
Annual Company Retreat