Site Reliability / Production Engineer @Storyteller
Software Development
Salary up to eur 32,00..
Remote Location
Employment Type full-time
Posted 2mths ago

[Hiring] Site Reliability / Production Engineer @Storyteller

2mths ago - Storyteller is hiring a remote Site Reliability / Production Engineer. πŸ’Έ Salary: up to eur 32,000 per year πŸ“Location: Turkey

Role Description

We are looking for hands-on production engineers who can take ownership when live systems need attention: establish the customer impact, investigate the evidence, take safe action and keep the response moving.

You will use AI throughout the work, but not as a substitute for judgement. You will be expected to supervise its output, understand the risk of any action and validate that the real customer outcome has recovered.

Responsibilities

  • Respond to live incidents
    • Receive automated alerts and technical escalations from Support.
    • Establish customer impact, severity, blast radius and the current system state.
    • Investigate using logs, metrics, traces, dashboards, deployment history, infrastructure, databases, queues, background jobs, APIs and application code.
    • Use AI throughout triage and diagnosis while checking its conclusions against real evidence.
    • Choose and execute a proportionate mitigation, rollback, repair or bounded fix.
    • Validate that the customer outcome has recovered.
    • Keep ownership, uncertainty, decisions and next actions visible.
    • Join customer conversations occasionally when direct technical involvement is genuinely useful.
  • Coordinate the right response
    • Bring in the relevant product team when an incident requires deep product knowledge.
    • Escalate with evidence, customer impact, actions already taken and the specific decision or help required.
    • Protect developers from routine pages.
    • Produce a clear incident record and handover.
  • Improve the reliability system
    • Remove, consolidate and tune low-value alerts.
    • Analyse material incidents with AI and validate the conclusions.
    • Improve dashboards, diagnostics, service ownership and escalation information.
    • Create safe, supervised automation for common operational actions.
    • Work with product teams to close observability, rollback, runbook and supportability gaps.
    • Detect and help contain unusual service-cost behaviour.
    • Make reliability and on-call performance easier for the company to understand and improve over time.

Qualifications

  • Agency and ownership: You take responsibility for ambiguous live problems.
  • Operational judgement: You can separate customer impact, symptoms and likely causes.
  • Technical comfort and aptitude: You are comfortable exploring unfamiliar systems.
  • AI-native execution: You use AI for substantive technical work.
  • Accuracy and validation discipline: You actively look for false confidence.
  • Systems thinking: You look for repeated patterns and improve the triggers.
  • Clear coordination and communication: You communicate calmly and concisely.
  • Curiosity and resilience: You learn unfamiliar products and tools quickly.

Nice to have

  • Cloud platforms such as Azure or Cloudflare.
  • Distributed application and API diagnostics.
  • Databases, queues and background-processing systems.
  • Observability, alerting and incident-management platforms.
  • Infrastructure, deployment and release automation.
  • Application development and safe production debugging.
  • AI coding agents and workflow automation.

Recruitment Process

  • Hiring Manager Conversation (20-30 mins): A short call to get to know you.
  • Paid Take-home Task (~60-90 mins): A small, bounded production-incident exercise.
  • Task Review and CTO Interview (60-75 mins): Review your submission and discuss decisions.

Working Pattern

  • This role provides out-of-hours production coverage.
  • The two hires will share an agreed rota.
  • Active shifts will be divided between both hires with appropriate rest days.
  • Weekday pager coverage from 01:00-06:00 UK time.
  • When there are no live incidents, the active shift will be used for reliability-improvement work.
Before You Apply
️
remote Be aware of the location restriction for this remote position: Turkey
β€Ό Beware of scams! When applying for jobs, you should NEVER have to pay anything. Learn more.
Site Reliability / Production Engineer @Storyteller
Software Development
Salary up to eur 32,00..
Remote Location
Employment Type full-time
Posted 2mths ago
Apply for this position
Did not apply βœ“
Applied βœ“
Sent Follow-Up βœ“
Interview Scheduled βœ“
Interview Completed βœ“
Offer Accepted βœ“
Offer Declined βœ“
Application Denied βœ“
Unlock 125,000+ Remote Jobs
️
remote Be aware of the location restriction for this remote position: Turkey
β€Ό Beware of scams! When applying for jobs, you should NEVER have to pay anything. Learn more.
Apply for this position
Did not apply βœ“
Applied βœ“
Sent Follow-Up βœ“
Interview Scheduled βœ“
Interview Completed βœ“
Offer Accepted βœ“
Offer Declined βœ“
Application Denied βœ“
Unlock 125,000+ Remote Jobs
Γ—
Apply to the best remote jobs
before everyone else

Access 125,000+ vetted remote jobs and get daily alerts.

4.9 β˜…β˜…β˜…β˜…β˜… from 500+ reviews

⚑ 127,023+ remote jobs, refreshed hourly

πŸ”” Real-time alerts: Apply first, direct to employer

πŸ›‘οΈ Vetted companies, no scams, true remote only

Unlock All Jobs Now

Maybe later