Senior Site Reliability Engineer @Lodgify
Software Development
Salary unspecified
Remote Location
Employment Type full-time
Posted 1mth ago

[Hiring] Senior Site Reliability Engineer @Lodgify

1mth ago - Lodgify is hiring a remote Senior Site Reliability Engineer. πŸ’Έ Salary: unspecified πŸ“Location: Spain

Role Description

Are you a systems-minded engineer who cares deeply about reliability, scalability, and production excellence? Join Lodgify as a Senior Site Reliability Engineer and help our engineering teams build and operate services that are reliable, observable, scalable, and resilient by design. In this role, you will improve the reliability of our shared infrastructure and product services while helping teams own their systems in production. You will work in the Platform team to strengthen observability, reduce operational toil, improve incident response, define practical SRE standards, and improve the reliability of critical infrastructure and delivery workflows.

  • Define meaningful SLIs, SLOs, and reliability targets for the platform.
  • Collaborate with the software engineering teams to define and achieve the best practices for software observability, SLIs, SLOs and reliability.
  • Strengthen production readiness by improving service ownership, observability, alerting, runbooks, scaling assumptions, rollback paths, and failure-mode preparedness.
  • Improve the reliability, scalability, and performance of cloud, Kubernetes, and shared infrastructure.
  • Build actionable observability using metrics, logs, traces, and golden signals, with tools such as Datadog, Prometheus, and Grafana.
  • Implement operational and security best practices through guidelines, policies and automation.
  • Reduce alert noise and improve signal quality so teams can detect, understand, and resolve issues quickly.
  • Automate repetitive operational work using Python or other languages.
  • Implement self-service Internal Developer Platform features via APIs and Kubernetes operators.
  • Improve deployment safety, rollbackability, and release observability.
  • Improve reliability of critical stateful systems such as databases, caches, queues, and streaming platforms.
  • Participate in on-call, troubleshoot, and coordinate incident response.
  • Execute disaster recovery drills and analyse cloud/platform usage to identify cost and resource-efficiency gains.

Qualifications

  • 7+ years of production experience operating Kubernetes-based platforms and cloud infrastructure.
  • Understanding and application of SRE practices: SLIs, SLOs, error budgets, production readiness, incident response, post-incident learning, toil reduction, scalability, capacity planning, high availability, backups, and disaster recovery.
  • Ability to design and improve observability and alerting for critical systems.
  • Experience with stateful production systems such as relational databases, caches, queues, or streaming platforms.
  • Comfortable working in a transitional environment where SRE practices are being introduced.
  • Effective collaboration with Engineering, Platform, Security, and Product stakeholders.
  • Clear communication, good documentation, and enjoyment in coaching teams.
  • Model initiative and accountability, raising risks early and driving improvements through to completion.

Requirements

  • Critical services have clear owners, meaningful SLIs/SLOs, actionable alerts, dashboards, runbooks, and production readiness coverage.
  • Reliability targets are consistently met across critical infrastructure and services.
  • Operational toil and manual intervention are measurably reduced through automation and safer workflows.
  • MTTR improves through reduced alert noise, better signal quality, stronger observability, and clear incident response playbooks and escalation paths.
  • Post-incident actions are tracked, completed, and used to reduce repeat incidents.
  • Disaster recovery exercises validate that critical services and infrastructure can recover within agreed expectations.
  • Cloud and infrastructure resources are optimised without sacrificing performance, elasticity, or resilience.

Benefits

  • Remote Flexibility: The freedom to work from home any day that works for you.
  • Time to Recharge: 25 working days of paid vacation and Jornada Intensiva in August.
  • Alan Health Insurance: Premium health, dental, and mental health support via Alan. Pre-existing conditions are covered.
  • Meal Perk: €150/month allowance on your Alan card + 50% off Ametller Origen prepared dishes at the office.
  • Tax-Free Savings: Increase your take-home pay by using Flexible Remuneration for extra meal costs (up to €70/mo) and public transport (up to €136/mo).
  • Home Office Gear: We provide a table, ergonomic chair, and monitor for your home setup.
  • Language Learning: Free Spanish classes.
  • Referrals: Cash rewards for bringing in new talent.
  • Social Life: Daily office breakfast and monthly team events.
  • Dynamic Hub: A high-energy, inclusive environment designed for collaboration and connection.
Before You Apply
️
remote Be aware of the location restriction for this remote position: Spain
β€Ό Beware of scams! When applying for jobs, you should NEVER have to pay anything. Learn more.
Senior Site Reliability Engineer @Lodgify
Software Development
Salary unspecified
Remote Location
Employment Type full-time
Posted 1mth ago
Apply for this position
Did not apply βœ“
Applied βœ“
Sent Follow-Up βœ“
Interview Scheduled βœ“
Interview Completed βœ“
Offer Accepted βœ“
Offer Declined βœ“
Application Denied βœ“
Unlock 125,000+ Remote Jobs
️
remote Be aware of the location restriction for this remote position: Spain
β€Ό Beware of scams! When applying for jobs, you should NEVER have to pay anything. Learn more.
Apply for this position
Did not apply βœ“
Applied βœ“
Sent Follow-Up βœ“
Interview Scheduled βœ“
Interview Completed βœ“
Offer Accepted βœ“
Offer Declined βœ“
Application Denied βœ“
Unlock 125,000+ Remote Jobs
Γ—
Apply to the best remote jobs
before everyone else

Access 125,000+ vetted remote jobs and get daily alerts.

4.9 β˜…β˜…β˜…β˜…β˜… from 500+ reviews

⚑ 126,845+ remote jobs, refreshed hourly

πŸ”” Real-time alerts: Apply first, direct to employer

πŸ›‘οΈ Vetted companies, no scams, true remote only

Unlock All Jobs Now

Maybe later