Senior Site Reliability Engineer @Blackpoint Cyber
Software Development
Salary cad 131,000 - 1..
Remote Location
Employment Type full-time
Posted 1mth ago

[Hiring] Senior Site Reliability Engineer @Blackpoint Cyber

1mth ago - Blackpoint Cyber is hiring a remote Senior Site Reliability Engineer. πŸ’Έ Salary: cad 131,000 - 164,250 per year πŸ“Location: Canada

Role Description

We're hiring a Senior Site Reliability Engineer to design, implement, and maintain our cloud and on-premise infrastructure and CI/CD pipelines, with a focus on automation, scalability, and performance. You'll work across cloud platform administration, container orchestration, data streaming, observability, and incident response β€” partnering with engineering teams to keep our systems reliable, secure, and efficient, and helping foster a culture of continuous improvement.

Responsibilities

  • Design, develop, and maintain highly scalable infrastructure using Infrastructure as Code (Terraform and Terragrunt) for automated cloud resource provisioning and orchestration.
  • Own and optimize our AWS cloud environment, ensuring cost efficiency, security best practices, and high-availability standards.
  • Manage and optimize Kubernetes cluster environments (Helm, ArgoCD, Istio, Kustomize) to support continuous delivery and infrastructure-as-code practices.
  • Administer and scale data streaming infrastructure (Confluent Cloud, Apache Kafka) to support enterprise-level data processing.
  • Deploy, configure, and maintain Redis for caching and real-time data processing.
  • Implement and maintain monitoring, alerting, and incident response frameworks (Prometheus, Grafana, Alert Manager, OpsGenie/PagerDuty) to ensure system reliability and performance.
  • Facilitate controlled feature deployments and progressive rollouts through LaunchDarkly/PostHog.
  • Partner with software development teams to ensure seamless integration of new services, applications, and features into existing infrastructure.
  • Diagnose and resolve complex system-level issues, implementing solutions that maintain high performance and maximize uptime.
  • Drive continuous improvement of automation tooling, operational processes, and engineering methodologies to enhance scalability, reliability, and maintainability.
  • Stay current on emerging SRE trends and tools, and help the team adopt relevant industry advancements and best practices.

Qualifications

  • 5+ years of experience in a Senior Site Reliability Engineer role or equivalent, with substantial emphasis on cloud infrastructure management and automation.
  • Expertise in Infrastructure as Code (Terraform, Terragrunt) for enterprise-scale deployments.
  • Comprehensive knowledge of AWS, including designing, implementing, and maintaining secure, scalable, resilient cloud architectures.
  • Extensive hands-on experience with distributed data streaming (Confluent Cloud, Apache Kafka).
  • Proven experience with Redis for caching and Amazon RDS for relational database management.
  • Experience with enterprise search and analytics platforms (OpenSearch, Elasticsearch, ChaosSearch).
  • Proficiency designing and implementing monitoring/alerting infrastructure (Prometheus, Grafana, Alert Manager, OpsGenie/PagerDuty).
  • Practical experience with feature flag systems (LaunchDarkly/PostHog) for controlled release management.
  • Extensive experience administering production-grade Kubernetes (Helm, ArgoCD, Istio); working knowledge of Kustomize.
  • Strong problem-solving skills, with the ability to troubleshoot complex systems in production.
  • Strong communication and collaboration skills, with experience working in Agile environments.

Requirements

  • Multi-cloud experience (Google Cloud Platform, Microsoft Azure).
  • Understanding of security frameworks and compliance standards for cloud-native/containerized environments.
  • Serverless computing and CI/CD pipeline experience (Jenkins, GitHub Actions).
  • Software development proficiency in Node.js, Python, and/or Go.

Benefits

  • Competitive Health, Vision, Dental, and Life Insurance plans for eligible employees in the US.
  • Robust 401k plan.
  • Discretionary Time Off and other minor perks.
  • Competitive benefits for international employees in accordance with local market standards and applicable country requirements.
  • Equity participation available to employees globally, with program details varying by location and employment structure.
Before You Apply
️
remote Be aware of the location restriction for this remote position: Canada
β€Ό Beware of scams! When applying for jobs, you should NEVER have to pay anything. Learn more.
Senior Site Reliability Engineer @Blackpoint Cyber
Software Development
Salary cad 131,000 - 1..
Remote Location
Employment Type full-time
Posted 1mth ago
Apply for this position
Did not apply βœ“
Applied βœ“
Sent Follow-Up βœ“
Interview Scheduled βœ“
Interview Completed βœ“
Offer Accepted βœ“
Offer Declined βœ“
Application Denied βœ“
Unlock 125,000+ Remote Jobs
️
remote Be aware of the location restriction for this remote position: Canada
β€Ό Beware of scams! When applying for jobs, you should NEVER have to pay anything. Learn more.
Apply for this position
Did not apply βœ“
Applied βœ“
Sent Follow-Up βœ“
Interview Scheduled βœ“
Interview Completed βœ“
Offer Accepted βœ“
Offer Declined βœ“
Application Denied βœ“
Unlock 125,000+ Remote Jobs
Γ—
Apply to the best remote jobs
before everyone else

Access 125,000+ vetted remote jobs and get daily alerts.

4.9 β˜…β˜…β˜…β˜…β˜… from 500+ reviews

⚑ 127,023+ remote jobs, refreshed hourly

πŸ”” Real-time alerts: Apply first, direct to employer

πŸ›‘οΈ Vetted companies, no scams, true remote only

Unlock All Jobs Now

Maybe later