Role Description
As the Senior Engineering Manager of Site Reliability Engineering, you will lead a team responsible for improving the reliability and operational maturity of Upstart’s products and services. You will drive high impact improvements across incident management, observability, operational readiness, and reliability engineering.
-
Serve as the accountable leader of the SRE function, translating reliability strategy into focused plans, clear ownership, and measurable outcomes.
-
Partner closely with engineering leaders to establish reliability expectations, identify systemic risks, and build scalable capabilities that enable teams to operate services safely and independently.
-
Balance immediate operational needs with durable improvements that reduce risk, strengthen resilience, and improve how Upstart learns from production.
How you’ll make an impact
-
Team Leadership and Execution
-
Manage and develop a team focused on incident management, observability, operational readiness, and reliability engineering.
-
Define a clear charter, priorities, roadmap, and measurable outcomes for the SRE function.
-
Translate strategy into capacity aware plans with explicit trade offs, ownership, milestones, and success measures.
-
Maintain visibility into delivery health, operational risks, and team performance, intervening early when execution drifts.
-
Build a resilient operating model through cross-training, shared context, effective delegation, and clear primary and secondary ownership.
-
Set a high bar for technical quality, operating rigor, and executive communication.
-
Develop engineers and leaders who can independently own complex reliability initiatives.
-
Incident Management and Learning
-
Evolve Upstart’s incident management program to improve detection, response, coordination, communication, and recovery.
-
Establish clear standards for managing high severity incidents and provide visible leadership during critical events.
-
Improve postmortem quality and ensure incident learnings result in durable engineering improvements.
-
Identify recurring failure patterns and drive systemic solutions across teams.
-
Create strong feedback loops from incidents into roadmaps, service standards, operational readiness requirements, and measurable risk reduction.
-
Observability and Reliability Engineering
-
Improve the quality, accessibility, and trustworthiness of signals used to understand production health.
-
Drive consistent practices across metrics, logs, traces, alerting, and service health.
-
Advance the use of service level objectives and customer impact signals to guide priorities and operational decisions.
-
Reduce detection gaps, noisy alerts, manual investigation, and recurring operational toil.
-
Define measurable reliability outcomes and use data to prioritize investments and communicate impact.
-
Partner with platform and product engineering teams to embed reliability into standard engineering workflows.
-
Operational Readiness and Resilience
-
Establish scalable operational readiness standards for new services, major launches, and architectural changes.
-
Set clear expectations for service ownership, monitoring, capacity, failure handling, and incident response.
-
Identify systemic reliability risks and partner with engineering teams to prioritize and address them.
-
Improve resilience through automation, failure testing, recovery capabilities, and operational safeguards.
-
Build operating mechanisms that turn reviews and analysis into clear decisions, owners, timelines, and sustained follow through.
-
Align stakeholders and dependencies before critical launches and engineering decisions.
Qualifications
-
5+ years of reliability engineering management experience and 7+ years of experience in software engineering, site reliability engineering, infrastructure, or platform engineering.
-
Significant hands-on experience in Site Reliability Engineering, Production Engineering, or an equivalent role responsible for operating and improving production systems.
-
Direct experience managing an SRE, Production Engineering, or equivalent reliability function, including ownership of its strategy, roadmap, operating model, and outcomes.
-
Strong technical depth in distributed systems, cloud infrastructure, observability, and production operations.
-
Experience leading high severity incident response and improving incident management practices at scale.
-
Demonstrated ability to translate strategy into focused, capacity aware plans and deliver measurable outcomes.
-
Track record of hiring, developing, and retaining high performing engineers and engineering leaders.
-
Strong cross-functional leadership and communication, with the ability to turn complex operational data into clear decisions and drive alignment across teams.
Preferred Qualifications
-
Experience operating large scale, highly available distributed systems.
-
Experience implementing or evolving service-level objectives and error-budget practices.
-
Experience with observability platforms such as Datadog, Grafana, Prometheus, OpenTelemetry, or similar technologies.
-
Experience developing incident management, operational readiness, or resilience programs across a large engineering organization.
-
Familiarity with Kubernetes, AWS, and modern cloud native architectures.
-
Experience supporting major platform or architectural transitions.
-
Strong product mindset when building internal reliability capabilities.
-
Experience establishing executive level reliability reporting and operating reviews.
Benefits
-
Competitive compensation, including base pay, bonus opportunities, and annual equity grants that vest quarterly.
-
Retirement benefits to help you plan for the future, including a 401(k) or Group Retirement Savings Plan with a company match of $2 for every $1 contributed, up to $15,000 annually (USD in the US, CAD in Canada).
-
Employee Stock Purchase Plan (ESPP) with discounted stock purchase options for eligible employees (US only).
-
Comprehensive health coverage designed to support you and your family, including medical, dental, vision, and wellness resources for US and supplemental health coverage for Canada.
-
Health Savings Account contributions from Upstart for eligible plans (US only).
-
Income protection benefits, including life insurance and disability coverage for added financial security.
-
Paid time off, sick leave, and company holidays, in line with local requirements.
-
Paid family and parental leave to support caregiving and major life moments (duration varies by country).
-
Family-centered benefits to support fertility, parenthood, and caregiving needs.
-
Employee Assistance Program (EAP) offering mental health support and life-centered resources.
-
Financial wellness resources, including access to financial planning tools and a financial concierge service (US Only).
-
Annual wellness allowance to support your physical and emotional well-being and personal development, based on what matters most to you.
-
Annual productivity allowance to invest in relevant tools and resources you need to do your best work, no matter where you work from.
-
Connection and community through team events, all-company updates, and employee resource groups (ERGs).
-
Onsite perks, including catered lunches and fully stocked micro-kitchens when working from one of our offices in the Bay Area, Austin, Columbus, and New York City (opening Summer 2026!).