Site Reliability Engineer @GiveCampus
All Others
Salary unspecified
Remote Location
πŸ‡ΊπŸ‡Έ USA Only
Employment Type full-time
Posted 2wks ago

[Hiring] Site Reliability Engineer @GiveCampus

2wks ago - GiveCampus is hiring a remote Site Reliability Engineer. πŸ’Έ Salary: unspecified πŸ“Location: USA

Role Description

GiveCampus is looking for a hands-on Site Reliability Engineer to help improve the reliability, performance, and operational maturity of our platform. With our migration to AWS complete, this role will focus on operating and strengthening our production environment:

  • Improving observability
  • Automating infrastructure and operational work
  • Responding to incidents
  • Partnering with product engineers to build resilient systems

You will own well-scoped reliability projects and contribute to larger cross-functional initiatives. You should be comfortable working independently on straightforward problems, asking for guidance when needed, explaining technical tradeoffs, and keeping teammates informed at important milestones.

What you'll do:

  • Operate, maintain, and improve production infrastructure in AWS.
  • Build and maintain infrastructure as code using Terraform.
  • Support workloads running on Kubernetes and Amazon EKS.
  • Improve dashboards, alerts, and service-level indicators using New Relic or comparable observability platforms.
  • Investigate production issues, identify root causes, and implement durable fixes.
  • Participate in the shared 24/7 on-call rotation and contribute to effective incident response.
  • Participate in blameless postmortems and complete follow-up actions that reduce the likelihood or impact of repeat incidents.
  • Partner with product engineers to troubleshoot performance and reliability issues throughout the application stack.
  • Improve application resilience using established patterns such as timeouts, retries, queuing, backpressure, and idempotency.
  • Maintain and improve CI/CD pipelines and deployment workflows using tools such as GitHub Actions and CircleCI.
  • Automate repetitive operational tasks and identify opportunities to reduce engineering toil.
  • Create and maintain runbooks, system diagrams, troubleshooting guides, and production documentation.
  • Contribute to capacity planning, performance testing, database reliability, and production-readiness reviews.
  • Apply established security, access-control, logging, and compliance practices to infrastructure work.
  • Own small-to-medium reliability improvements from technical design and work breakdown through delivery.
  • Communicate progress, risks, tradeoffs, and blockers clearly while incorporating feedback from engineering partners.

Qualifications

  • Approximately 5+ years of related experience in software engineering, infrastructure, systems engineering, SRE, Platform Engineering, DevOps, or equivalent practical experience.
  • Hands-on experience operating production workloads in AWS.
  • Experience building or maintaining infrastructure using Terraform or a similar infrastructure-as-code tool.
  • Experience with New Relic, Datadog, or another modern observability platform.
  • Experience troubleshooting production incidents and participating in an on-call rotation.
  • Experience building or maintaining CI/CD pipelines.
  • Software development or scripting experience, with the ability to read, debug, and make targeted changes to application or automation code.
  • Working knowledge of Linux, networking, distributed systems, and relational databases.
  • Ability to articulate root causes, explain technical tradeoffs, and translate findings into practical solutions.
  • Ability to manage a well-scoped project with general direction and provide timely updates at key milestones.
  • Strong written and verbal communication skills and a collaborative approach to working across engineering disciplines.
  • A habit of automating repetitive work and improving the reliability of the systems you support.

Bonus points

  • Experience with Ruby or Ruby on Rails.
  • PostgreSQL administration or performance-tuning experience.
  • Experience with Kubernetes and Amazon EKS.
  • Experience with Redis, OpenSearch, or Amazon RDS.
  • Experience operating enterprise SaaS products at scale.
  • Familiarity with SLOs, SLIs, error budgets, capacity modeling, or load testing.
  • Experience with payments, fintech, or other regulated systems.
  • Experience supporting SOC 2 or similar security and compliance programs.

Company Description

At GiveCampus, we value diversity and we pledge to foster an environment of support, inclusivity, and learning, both on the job and throughout the application process. In this spirit, we encourage candidates of all backgrounds to apply.

GiveCampus is an Equal Opportunity Employer. Applicants and employees are not discriminated against because of race, color, creed, sex, sexual orientation, gender identity or expression, age, religion, national origin, citizenship status, disability, ancestry, marital status, veteran status, medical condition or any protected category prohibited by local, state or federal laws.

If you feel like you don't meet all of the requirements for this role, please apply anyways. We know confidence gaps and imposter syndrome often get in the way of connecting with incredible people, and we don't want them to prevent us from meeting you.

Before You Apply
️
πŸ‡ΊπŸ‡Έ Be aware of the location restriction for this remote position: USA Only
β€Ό Beware of scams! When applying for jobs, you should NEVER have to pay anything. Learn more.
Site Reliability Engineer @GiveCampus
All Others
Salary unspecified
Remote Location
πŸ‡ΊπŸ‡Έ USA Only
Employment Type full-time
Posted 2wks ago
Apply for this position
Did not apply βœ“
Applied βœ“
Sent Follow-Up βœ“
Interview Scheduled βœ“
Interview Completed βœ“
Offer Accepted βœ“
Offer Declined βœ“
Application Denied βœ“
Unlock 125,000+ Remote Jobs
️
πŸ‡ΊπŸ‡Έ Be aware of the location restriction for this remote position: USA Only
β€Ό Beware of scams! When applying for jobs, you should NEVER have to pay anything. Learn more.
Apply for this position
Did not apply βœ“
Applied βœ“
Sent Follow-Up βœ“
Interview Scheduled βœ“
Interview Completed βœ“
Offer Accepted βœ“
Offer Declined βœ“
Application Denied βœ“
Unlock 125,000+ Remote Jobs
Γ—
Apply to the best remote jobs
before everyone else

Access 125,000+ vetted remote jobs and get daily alerts.

4.9 β˜…β˜…β˜…β˜…β˜… from 500+ reviews

⚑ 127,028+ remote jobs, refreshed hourly

πŸ”” Real-time alerts: Apply first, direct to employer

πŸ›‘οΈ Vetted companies, no scams, true remote only

Unlock All Jobs Now

Maybe later