Staff Software Engineer, Reliability @CommandLink
Software Development
Salary unspecified
Remote Location
Employment Type full-time
Posted 3wks ago

[Hiring] Staff Software Engineer, Reliability @CommandLink

3wks ago - CommandLink is hiring a remote Staff Software Engineer, Reliability. πŸ’Έ Salary: unspecified πŸ“Location: Latin America (LATAM)

Role Description

Command|Alert is CommandLink's signal-processing core, the engine that turns raw security, monitoring, and customer-defined telemetry into alerts customers actually trust. Alert fatigue and noise are the top complaint across every competitor in this space, and this role exists to make sure our alerts are the ones people don't tune out.

As a Staff Software Engineer on Command|Alert, you'll own the reliability of the alerting pipeline end to end and drive the org's most consequential decisions on how it's architected. This role suits someone who thinks like a site reliability engineer as much as a software engineer: fluent in SLOs, error budgets, and blameless incident response, with the technical range to reason across security tooling, monitoring telemetry, syslog, OpenTelemetry, and L2-L4 network protocols.

Key Responsibilities:

  • Own the reliability of the alerting pipeline end to end, from OpenSearch alert evaluation through Kafka delivery via OpenSearch callbacks to downstream notification, including idempotency guarantees and soak-tested behavior under sustained load.
  • Define SLOs and SLIs for the pipeline's availability, latency, and delivery guarantees, and use error budgets to guide how much investment goes into reliability work versus new capability.
  • Lead incident response for the pipeline's most critical production issues, running blameless post-mortems and driving the systemic fixes and automation that come out of them.
  • Build the observability and automation needed to run the pipeline at Fortune 1000 scale with low operational toil, and own capacity planning as ingestion volume and customer count grow.
  • Set the architecture for how Command|Alert evaluates rule-based thresholds, ML anomaly scores, and correlation logic, turning diverse telemetry into usable network and system topologies that power LLM-driven investigation and remediation.
  • Mentor engineers across the teams you touch and represent Command|Alert's technical direction to stakeholders outside engineering.
  • Takes on additional responsibilities and projects as needed to support the success of the team and organization.

Qualifications

  • Background operating as a Site Reliability Engineer, DevOps engineer, or in a similar production-ownership role, with fluency in SLOs, SLIs, error budgets, and blameless incident response.
  • Demonstrated experience designing, building, or operating high-reliability alerting or notification systems in production, including rule-based and ML-based detection at scale.
  • Strong Kafka experience, and a track record building systems where webhook reliability, idempotency, and delivery guarantees under load are non-negotiable.
  • Experience building the observability and automation that let a high-volume production system run with low operational toil.
  • A working command of telemetry and protocol data (security tooling output, syslog, OpenTelemetry, NetFlow/sFlow, SNMP, ICMP, firewall logs) and the ability to turn it into real network and system topologies.
  • Recognized mastery of Go and/or Python, with range across container orchestration and a multi-cloud footprint, and a demonstrated ability to make org-level architecture calls.

Requirements

  • Experience with chaos engineering or fault-injection testing (e.g., Gremlin, Chaos Mesh).
  • Familiarity with SLO and error-budget tooling and practice (e.g., Nobl9, Google's SRE workbook approach).
  • Experience with Temporal or a comparable workflow orchestration platform.
  • Experience operating in a multi-tenant, cloud-native environment with secrets management and TLS at scale.

Benefits

  • Room to grow at a high-growth company.
  • An environment that celebrates ideas and innovation.
  • Your work will have a tangible impact.
  • Flexible time off.
  • Fun events at cool locations.
  • Employee referral bonuses to encourage the addition of great new people to the team.
Before You Apply
️
remote Be aware of the location restriction for this remote position: Latin America (LATAM)
β€Ό Beware of scams! When applying for jobs, you should NEVER have to pay anything. Learn more.
Staff Software Engineer, Reliability @CommandLink
Software Development
Salary unspecified
Remote Location
Employment Type full-time
Posted 3wks ago
Apply for this position
Did not apply βœ“
Applied βœ“
Sent Follow-Up βœ“
Interview Scheduled βœ“
Interview Completed βœ“
Offer Accepted βœ“
Offer Declined βœ“
Application Denied βœ“
Unlock 125,000+ Remote Jobs
️
remote Be aware of the location restriction for this remote position: Latin America (LATAM)
β€Ό Beware of scams! When applying for jobs, you should NEVER have to pay anything. Learn more.
Apply for this position
Did not apply βœ“
Applied βœ“
Sent Follow-Up βœ“
Interview Scheduled βœ“
Interview Completed βœ“
Offer Accepted βœ“
Offer Declined βœ“
Application Denied βœ“
Unlock 125,000+ Remote Jobs
Γ—
Apply to the best remote jobs
before everyone else

Access 125,000+ vetted remote jobs and get daily alerts.

4.9 β˜…β˜…β˜…β˜…β˜… from 500+ reviews

⚑ 126,920+ remote jobs, refreshed hourly

πŸ”” Real-time alerts: Apply first, direct to employer

πŸ›‘οΈ Vetted companies, no scams, true remote only

Unlock All Jobs Now

Maybe later