Lead Site Reliability Engineer @Cloudlinux
All Others
Salary unspecified
Remote Location
Employment Type full-time
Posted 1mth ago

[Hiring] Lead Site Reliability Engineer @Cloudlinux

1mth ago - Cloudlinux is hiring a remote Lead Site Reliability Engineer. πŸ’Έ Salary: unspecified πŸ“Location: Worldwide

Role Description

Imunify360 is a multi-layer Linux server security suite β€” WAF, IDS/IPS, malware scanning and cleanup, proactive defence, patch management, reputation β€” running as an agent on hundreds of thousands of customer servers, backed by a cloud estate of scanning, correlation and signature-delivery services on our own bare metal.

Your job is to make that class of failure detectable in hours instead of months, across the whole product line, and to build the system that keeps it detectable as the product changes. This is a greenfield charter inside a brownfield estate. You are not inheriting an SRE team, an SLO framework or a paging culture. You are defining them, with the engineering leads, and then making them stick.

What you'll do

  • Define what "working" means for ~70 components
    • Run SLI definition with squad leads and senior engineers.
    • Build the taxonomy this product actually needs, which is broader than availability and latency:
      • Service SLIs β€” availability, latency, error rate for cloud-side services.
      • Fleet SLIs β€” heartbeat reachability, version and configuration convergence across the installed base.
      • Control-efficacy SLIs β€” the differentiator.
      • Delivery SLIs β€” artifact publish success, rule-to-fleet lead time, hotfix time-to-convergence.
      • Pipeline SLIs β€” ingest lag, verdict latency, queue age, backlog burn.
    • Enforce one non-negotiable design rule: an SLI must be measurable from outside the gate of the thing it measures.
    • Attach an SLO, an error budget and an owning squad to each.
  • Build the collection system
    • Design and build the pipeline that gets these indicators off the fleet and into a queryable store.
    • Extend agent-side and service-side instrumentation where the signal does not exist yet, in Python, Go and Rust.
    • Consolidate the current sprawl of dashboards, ad-hoc queries and reporting paths into a defensible set of instruments.
  • Build alerting and alert management
    • Symptom-based, SLO-anchored alerting with multi-window burn-rate semantics.
    • A three-tier taxonomy β€” page / ticket / dashboard.
    • Every alert ships with an owner, a runbook and a documented failure mode.
    • Alert hygiene as a standing practice.
  • Build escalation
    • Component β†’ owning squad ownership map, kept current, machine-readable.
    • Severity matrix, acknowledgement SLAs, follow-the-sun rota design across UTCβˆ’5 … UTC+8.
    • Incident command practice and blameless postmortems within 24 hours.
    • Design the escalation system so that squads carry their own pagers.

Qualifications

  • Substantial production-engineering or SRE experience.
  • Strong Python skills.
  • Deep practical grip on time-series and event telemetry at scale.
  • Distributed systems debugging on bare metal and long-lived hosts.
  • Configuration management and CI at production scale.
  • The judgement to design measurement for machines you do not own and cannot scrape.
  • Written communication that holds up async.

Requirements

  • Security product background β€” WAF, EDR, AV, vulnerability management.
  • Monitoring under audit: SOC 2 CC7.x, ISO 27001 A.8.16, NIST SP 800-137 continuous monitoring.
  • OpenTelemetry, eBPF, Sentry.
  • Cost- and cardinality-aware telemetry design.
  • Fluency with agentic development tooling.
  • Kubernetes experience.

Benefits

  • A strong focus on professional development with opportunities for learning and growth.
  • Interesting and challenging projects.
  • Mentor and other knowledge-exchange programs.
  • Fully remote work with flexible working hours.
  • Paid 24 days of vacation per year, 10 days of national holidays, and unlimited sick leaves.
  • Compensation for private medical insurance.
  • Co-working and gym/sports reimbursement.
  • The opportunity to receive a reward for the most innovative idea that the company can patent.
Before You Apply
️
worldwide Be aware of the location restriction for this remote position: Worldwide
β€Ό Beware of scams! When applying for jobs, you should NEVER have to pay anything. Learn more.
Lead Site Reliability Engineer @Cloudlinux
All Others
Salary unspecified
Remote Location
Employment Type full-time
Posted 1mth ago
Apply for this position
Did not apply βœ“
Applied βœ“
Sent Follow-Up βœ“
Interview Scheduled βœ“
Interview Completed βœ“
Offer Accepted βœ“
Offer Declined βœ“
Application Denied βœ“
Unlock 125,000+ Remote Jobs
️
worldwide Be aware of the location restriction for this remote position: Worldwide
β€Ό Beware of scams! When applying for jobs, you should NEVER have to pay anything. Learn more.
Apply for this position
Did not apply βœ“
Applied βœ“
Sent Follow-Up βœ“
Interview Scheduled βœ“
Interview Completed βœ“
Offer Accepted βœ“
Offer Declined βœ“
Application Denied βœ“
Unlock 125,000+ Remote Jobs
Γ—
Apply to the best remote jobs
before everyone else

Access 125,000+ vetted remote jobs and get daily alerts.

4.9 β˜…β˜…β˜…β˜…β˜… from 500+ reviews

⚑ 126,991+ remote jobs, refreshed hourly

πŸ”” Real-time alerts: Apply first, direct to employer

πŸ›‘οΈ Vetted companies, no scams, true remote only

Unlock All Jobs Now

Maybe later