Role Description
We are seeking an experienced Site Reliability Engineer to build reliable, observable systems that enable client teams to operate with confidence at scale. You'll work embedded within cross-functional teams—partnering with developers, platform engineers, and operations teams—to reduce toil, improve system reliability, and prevent incidents before they impact users.
-
Take ownership of system reliability across our full backend stack—from AWS infrastructure (ALB, ECS/Fargate, Aurora) to Python services.
-
Design and maintain observability systems, define SLAs and error budgets, build incident response processes, and drive improvements in uptime and performance.
-
Participate in 24/7 P0 on-call rotations with a 10-minute acknowledgment SLA.
-
Lead incident investigations and postmortems, helping teams shift from reactive firefighting to proactive reliability engineering.
-
Mentor and onboard additional SRE engineers as the team scales from 4 to 8-12 over the coming months.
-
Advise leadership on reliability strategy and help shape how organizations approach operational excellence.
Qualifications
-
6+ years hands-on experience in SRE, DevOps, or Platform Engineering at scale.
-
Strong proficiency in one cloud platform: Deep AWS expertise (ALB, ECS/Fargate, RDS Aurora, Lambda, IAM).
-
Production Python backend engineering experience; comfortable debugging and optimizing Python services in containers and Lambdas.
-
Experience building or bootstrapping SRE programs from scratch.
-
Hands-on experience with incident response, on-call rotations, postmortems, and SLOs.
-
Deep experience designing and implementing observability (metrics, logging, tracing, alerting).
-
Solid understanding of containerization (Kubernetes, Docker) and deployment automation.
-
Experience writing and maintaining infrastructure-as-code (Terraform, CloudFormation, Pulumi, etc.).
-
Hands-on experience with CI/CD pipelines and deployment strategies.
-
Ability to troubleshoot distributed systems and production issues with debugging complex failures.
-
Clear communication skills and the ability to explain reliability trade-offs to stakeholders.
-
Scripting proficiency (Python, Bash, Go); Linux/Unix administration fundamentals.
-
Version control (Git), collaboration workflows, and infrastructure automation patterns.
Requirements
-
Design observability systems (metrics, logging, tracing) and define SLOs, error budgets, and monitoring strategies aligned with business needs.
-
Own 24/7 P0 on-call rotation with 10-minute acknowledgment SLA; validate and escalate AI-generated incident reports to platform teams.
-
Establish reliability standards and SLA targets for backend services.
-
Mentor and onboard additional SRE engineers as team scales to 8-12.
-
Participate in incident response, troubleshooting, root-cause analysis, and postmortems.
-
Implement reliability improvements, capacity planning, and performance optimization.
-
Build self-healing automation and toil-reduction initiatives to minimize manual operational work.
-
Collaborate with teams to integrate reliability and observability into delivery.
-
Support modern workloads including AI applications and services; understand their operational and reliability requirements.
-
Create runbooks, architecture documentation, and troubleshooting guides to enable independence.
-
Review infrastructure changes and contribute to reliability standards and consistency.
-
Actively tune systems for latency, throughput, and resource efficiency based on observability data.
Team Collaboration
-
Availability for regular working sessions and discovery workshops with client teams.
-
Ability to work independently while collaborating with client & Modus leaders.
-
Overlap with client business hours daily is expected.
-
Reliable high-speed internet is a must.
-
Documenting solutions, sharing learnings with the team, contributing to internal wikis and runbooks.
-
Pair program and collaborate on complex reliability and operational problems.
-
Welcome code and architecture reviews; seek input on reliability designs.
-
Ability to coordinate with security, platform, and development teams on reliability and operations.
Bonus Skills
-
Experience designing or operating internal developer platforms.
-
Familiarity with chaos engineering, resilience testing, or failure scenario planning.
-
Cloud security and compliance experience (audit readiness, security hardening).
-
Experience with AI infrastructure, LLM serving, or agentic system operations.
-
Mentoring junior engineers or leading technical design discussions.
-
Cost optimization and cloud FinOps practices.
-
Experience with AI-generated incident reports and validation workflows.
-
Python performance tuning and backend optimization.
-
Building SRE programs in high-growth or fast-scaling environments.
You’ll Love
-
Learn from every incident and continuously improve systems.
-
Solve challenging reliability, observability, and operational problems.
-
Work with modern cloud platforms, Kubernetes, infrastructure-as-code, and observability.
-
Make a direct impact on system uptime, performance, and operational efficiency.
-
Collaborate with great teams and help shape enterprise reliability and SRE best practices.
-
Support teams operating AI applications and services reliably and safely at scale.