Site Reliability Engineer @NOVACARD
Devops
Salary unspecified
Employment Type full-time
Posted 3wks ago

[Hiring] Site Reliability Engineer @NOVACARD

3wks ago - NOVACARD is hiring a remote Site Reliability Engineer. πŸ’Έ Salary: unspecified πŸ“Location: Latin America (LATAM), Eastern Europe

Role Description

We’re looking for a Site Reliability Engineer (SRE) to ensure the stability, performance, and reliability of our critical production systems. You’ll work at the intersection of development and operations β€” building automation tools, improving observability, and preventing incidents before they occur.

Key Responsibilities

  • Ensure the stability, performance, and fault tolerance of production systems.
  • Develop and maintain infrastructure automation and observability tools.
  • Monitor system health, respond to incidents, and perform root cause analysis (RCA).
  • Collaborate with development teams to improve scalability and reliability of services.
  • Define and manage SLIs, SLOs, and Error Budgets.
  • Lead incident response: organize recovery, document RCA, and run blameless post-mortems.
  • Configure and administer Grafana and Zabbix, design insightful dashboards, and fine-tune alerting.
  • Integrate and monitor external vendor systems, collaborating with vendor technical support when needed.

Qualifications

  • Fluent Russian, English B1+ (comfortable with technical documentation).
  • 3+ years of experience as an SRE, DevOps, or Infrastructure Engineer.
  • Strong understanding of observability principles (metrics, logs, traces).
  • Hands-on experience with Grafana and Zabbix (administration, configuration, alert optimization).
  • Experience working with AWS and CI/CD tools.
  • Practical knowledge of SLI/SLO/Error Budget frameworks.
  • Experience leading and documenting incidents and post-mortems.
  • Scripting skills for automation (Python, Bash, or Go).
  • Solid understanding of distributed systems and networking fundamentals.

Nice to Have

  • Experience monitoring and supporting mobile applications.
  • Familiarity with Terraform, Prometheus, Loki, ELK, or similar tools.
  • Experience working with Kubernetes and containerized environments.

Benefits

  • Fully remote work format.
  • Official employment under the Russian Labor Code (for residents of Russia); contractor collaboration available for candidates from other countries.
  • Opportunity to work in an international team on a new digital product for the Mexican market.
  • A data-driven environment where your contributions have a real impact.
Before You Apply
️
remote Be aware of the location restriction for this remote position: Latin America (LATAM), Eastern Europe
β€Ό Beware of scams! When applying for jobs, you should NEVER have to pay anything. Learn more.
Site Reliability Engineer @NOVACARD
Devops
Salary unspecified
Employment Type full-time
Posted 3wks ago
Apply for this position
Did not apply βœ“
Applied βœ“
Sent Follow-Up βœ“
Interview Scheduled βœ“
Interview Completed βœ“
Offer Accepted βœ“
Offer Declined βœ“
Application Denied βœ“
Unlock 125,000+ Remote Jobs
️
remote Be aware of the location restriction for this remote position: Latin America (LATAM), Eastern Europe
β€Ό Beware of scams! When applying for jobs, you should NEVER have to pay anything. Learn more.
Apply for this position
Did not apply βœ“
Applied βœ“
Sent Follow-Up βœ“
Interview Scheduled βœ“
Interview Completed βœ“
Offer Accepted βœ“
Offer Declined βœ“
Application Denied βœ“
Unlock 125,000+ Remote Jobs
Γ—
Apply to the best remote jobs
before everyone else

Access 125,000+ vetted remote jobs and get daily alerts.

4.9 β˜…β˜…β˜…β˜…β˜… from 500+ reviews

⚑ 126,845+ remote jobs, refreshed hourly

πŸ”” Real-time alerts: Apply first, direct to employer

πŸ›‘οΈ Vetted companies, no scams, true remote only

Unlock All Jobs Now

Maybe later