Role Description
As a Senior Site Reliability Engineer at Upsun, you will lead the evolution of our cloud application platform from traditional cloud operations into a proactive, automation-driven SRE model. You will own critical engineering workstreams that enhance system reliability, scalability, and operational efficiency across multi-cloud environments. Partnering closely with engineering, product, and platform teams, you will embed reliability and performance into every stage of the software delivery lifecycle. In this role, you will anticipate architectural bottlenecks, drive infrastructure-as-code practices, and establish robust observability standards that ensure long-term system stability and uptime for our global users.
What to expect
-
Drive reliability & observability strategy:
Architect and elevate system monitoring, alerting, and logging using Prometheus, Grafana, and ELK Stack, establishing actionable SLIs/SLOs aligned with core business metrics.
-
Automate infrastructure & workflows:
Eliminate operational toil by designing and implementing resilient, automated solutions using IaC tools like Terraform and Ansible across AWS, GCP, and Azure.
-
Scale CI/CD & delivery pipelines:
Optimize pipeline architectures for fast, secure, and zero-downtime releases, ensuring infrastructure resilience during high-volume deployment cycles.
-
Lead incident response & post-mortems:
Guide high-priority incident triage, drive blameless post-mortem analysis, and implement preventative measures to continuously improve system resiliency.
-
Cross-functional leadership:
Partner with product and software engineering teams to incorporate SRE best practices into product roadmaps.
-
Champion technical innovation:
Proactively identify performance bottlenecks and evaluate emerging technologies (e.g., eBPF, container orchestration) to optimize platform stability and performance.
-
Time distribution:
Follow a 4-week rotation balancing engineering and operations to focus on reliability, automation, and scalability through hands-on troubleshooting and engineering innovation.
Qualifications
-
5+ years of experience in Site Reliability Engineering, Cloud Operations, or DevOps, with proven experience owning reliability for production platforms at scale.
-
Strong proficiency in Go or Python to build custom automation tools, custom controllers, or SRE platform components (beyond basic shell scripting).
-
Advanced hands-on knowledge of Linux operating system internals, kernel parameters, networking protocols, performance profiling, and system troubleshooting.
-
Deep expertise with cloud providers (AWS, GCP, Azure, or Openstack) with custom tooling built around cloud SDKs, and declarative infrastructure tools (e.g., Terraform) to manage distributed systems.
-
Proven ability to anticipate operational risks, make architectural trade-offs, and lead technical infrastructure initiatives with minimal guidance.
-
Outstanding cross-functional communication skills with a track record of building alignment, and fostering an inclusive engineering culture.
Requirements
-
Experience with custom-built orchestration, edge, storage, and operational tooling in a dynamic environment.
-
Experience with Docker and production Kubernetes cluster management or containerized deployment architectures.
-
Familiarity with Platform-as-a-Service (PaaS) architectures or developer-facing cloud platforms.
Benefits
-
Flexible PTO
-
Company stock options
-
Professional development budget
-
Office equipment budget
-
Wellness budget
-
Annual team gatherings
-
Internet reimbursement
-
Inclusive parental leave
-
Remote work travel program