Site Reliability Engineer - Observability Platform @Ford
Software Development
Salary usd 85,400 - 19..
Remote Location
πŸ‡ΊπŸ‡Έ USA Only
Employment Type full-time
Posted 2d ago

[Hiring] Site Reliability Engineer - Observability Platform @Ford

2d ago - Ford is hiring a remote Site Reliability Engineer - Observability Platform. πŸ’Έ Salary: usd 85,400 - 192,900 per year πŸ“Location: USA

Role Description

We're hiring an experienced Site Reliability Engineer to architect, extend, and scale our global observability platform. This role sits at the intersection of software and systems engineering, requiring strong skills in distributed systems design, infrastructure automation, and production operations to ensure high availability, scalability, and maintainability of our monitoring stack.

  • Design and implement scalable observability pipelines spanning metrics, logging, tracing, and alerting.
  • Define and operationalize Service Level Indicators (SLIs) and Service Level Objectives (SLOs), and establish error budgets to effectively drive maximum availability and uptime.
  • Build reusable infrastructure-as-code templates and frameworks to standardize observability instrumentation and onboarding.
  • Architect, design, and develop automation to improve the resilience, recoverability, availability, and scalability of supported applications.
  • Leverage experience to safely perform destructive testing to seek and discover vulnerabilities.
  • Develop tooling to improve reliability, quality, and time-to-market for software solutions.
  • Identify and reduce or eliminate toil via automation to maximize time spent on engineering and innovation.
  • Collaborate with development teams to design, build, and operate scalable and resilient software systems using cloud-native principles.
  • Proactively identify stability risks and work with engineering leadership to establish appropriate mitigation plans.
  • Regularly review key technical metrics such as transaction errors, logging, response times, caching strategies, conversion/bounce rates, capacity, and resource utilization.
  • Conduct performance analysis and optimization of new and in-production systems, measuring and optimizing performance to get ahead of customer needs and drive continuous innovation.
  • Solve complex architecture, design, and business problems by simplifying processes, optimizing systems, and removing bottlenecks.
  • Recognize, validate, and evangelize emerging technologies and architectures that align with business objectives.
  • Troubleshoot complex, distributed production systems and drive root-cause analysis for platform-level incidents.
  • Participate in incident response, support, recovery, and postmortem analysis.
  • Provide technical guidance and mentorship to other team members.
  • Continuously evaluate and integrate AI/ML capabilities to enhance anomaly detection, alerting precision, and performance insights.
  • Collaborate cross-functionally with engineering teams to embed observability best practices into system design and deployment workflows.

Qualifications

  • Bachelor’s Degree in Computer Science or equivalent experience.
  • 3+ years of experience in an SRE role.
  • 5+ years of programming experience with one or more of: Python, Go, Java/Scala, C, or C++.
  • 3+ years of experience building reusable infrastructure-as-code templates & frameworks in Terraform or ToFu.
  • 3+ years of experience with APM and monitoring tools such as Dynatrace, New Relic, ELK, Splunk, Prometheus, Sensu, Nagios, Kafka, or DataDog.
  • 3+ years of experience with J2EE, NoSQL/SQL datastores, Spring Boot, GCP/AWS/Azure, and Docker/Kubernetes in developing multi-tier applications.
  • Experience with RESTful APIs and microservices platforms.
  • Working knowledge of the TCP/IP stack, internet routing, and load balancing.
  • Strong proficiency with Google Cloud Platform and its library of services.
  • Experience with automated, test-driven development in CI/CD pipelines.
  • Thorough understanding of software development and agile methodologies.
  • Understanding of, and ability to implement, effective observability strategies to improve MTTD/MTTR (Mean Time to Detect/Resolve).

Benefits

  • Immediate medical, dental, vision and prescription drug coverage.
  • Flexible family care days, paid parental leave, new parent ramp-up programs, subsidized back-up childcare and more.
  • Family building benefits including adoption and surrogacy expense reimbursement, fertility treatments, and more.
  • Vehicle discount program for employees and family members and management leases.
  • Tuition assistance.
  • Established and active employee resource groups.
  • Paid time off for individual and team community service.
  • A generous schedule of paid holidays, including the week between Christmas and New Year’s Day.
  • Paid time off and the option to purchase additional vacation time.
Before You Apply
️
πŸ‡ΊπŸ‡Έ Be aware of the location restriction for this remote position: USA Only
β€Ό Beware of scams! When applying for jobs, you should NEVER have to pay anything. Learn more.
Site Reliability Engineer - Observability Platform @Ford
Software Development
Salary usd 85,400 - 19..
Remote Location
πŸ‡ΊπŸ‡Έ USA Only
Employment Type full-time
Posted 2d ago
Apply for this position
Did not apply βœ“
Applied βœ“
Sent Follow-Up βœ“
Interview Scheduled βœ“
Interview Completed βœ“
Offer Accepted βœ“
Offer Declined βœ“
Application Denied βœ“
Unlock 125,000+ Remote Jobs
️
πŸ‡ΊπŸ‡Έ Be aware of the location restriction for this remote position: USA Only
β€Ό Beware of scams! When applying for jobs, you should NEVER have to pay anything. Learn more.
Apply for this position
Did not apply βœ“
Applied βœ“
Sent Follow-Up βœ“
Interview Scheduled βœ“
Interview Completed βœ“
Offer Accepted βœ“
Offer Declined βœ“
Application Denied βœ“
Unlock 125,000+ Remote Jobs
Γ—
Apply to the best remote jobs
before everyone else

Access 125,000+ vetted remote jobs and get daily alerts.

4.9 β˜…β˜…β˜…β˜…β˜… from 500+ reviews

⚑ 129,737+ remote jobs, refreshed hourly

πŸ”” Real-time alerts: Apply first, direct to employer

πŸ›‘οΈ Vetted companies, no scams, true remote only

Unlock All Jobs Now

Maybe later