Site Reliability Engineering SRE Leader with AI Experience @ZENITH INFOTEK LLC
Devops
Salary unspecified
Remote Location
πŸ‡ΊπŸ‡Έ USA Only
Employment Type contract
Posted 2mths ago

[Hiring] Site Reliability Engineering SRE Leader with AI Experience @ZENITH INFOTEK LLC

2mths ago - ZENITH INFOTEK LLC is hiring a remote Site Reliability Engineering SRE Leader with AI Experience. πŸ’Έ Salary: unspecified πŸ“Location: USA

Role Description

We are seeking an experienced Site Reliability Engineering (SRE) Leader with strong expertise in Java application development, cloud technologies, DevOps, and AI-driven solutions. The ideal candidate will lead initiatives focused on improving application reliability, scalability, automation, and operational excellence while leveraging modern AI and Generative AI technologies. This role requires close collaboration with engineering, infrastructure, and operations teams to build highly available, resilient, and high-performing enterprise applications.

Key Responsibilities

  • Design, develop, and implement scalable, reliable, and high-performance enterprise applications.
  • Collaborate with development, infrastructure, and operations teams to improve system reliability and availability.
  • Build and enhance automation solutions to streamline deployment, monitoring, and operational processes.
  • Develop and maintain CI/CD pipelines using Jenkins, CloudBees, or similar tools.
  • Monitor application health using observability platforms such as New Relic, Datadog, Dynatrace, and Splunk.
  • Analyze system performance, identify bottlenecks, and implement performance optimization strategies.
  • Design solutions for high availability, redundancy, auto-recovery, and fault tolerance.
  • Perform capacity planning and manage application performance against SLA/SLO objectives.
  • Lead Proof of Concepts (POCs) and successfully scale them into enterprise-wide implementations.
  • Implement best practices for application lifecycle management, monitoring, and continuous improvement.
  • Work with AWS cloud technologies to deploy and manage cloud-native applications.
  • Evaluate and integrate AI and Generative AI technologies into operational workflows.
  • Mentor engineering teams on SRE principles, DevOps practices, and reliability engineering.
  • Stay current with emerging technologies and recommend innovative solutions to improve platform stability and efficiency.

Qualifications

  • Bachelor's degree in Computer Science, Information Technology, Engineering, or a related field.
  • 10+ years of experience in enterprise Java application development and runtime support.
  • Strong experience with:
    • Java
    • Spring Boot
    • Microservices Architecture
    • Java Enterprise Frameworks
  • Deep understanding of:
    • Site Reliability Engineering (SRE)
    • Reliability Engineering concepts
    • Automation
    • High Availability
    • Auto Recovery
    • Scalability
    • Redundancy
  • Experience with AWS Cloud services (AWS Certification preferred).
  • Hands-on experience with CI/CD tools such as Jenkins or CloudBees.
  • Strong knowledge of observability and monitoring tools including:
    • New Relic
    • Datadog
    • Dynatrace
    • Splunk
  • Experience with PCF (Pivotal Cloud Foundry), App Pilot, and Presto.
  • Strong understanding of application infrastructure, runtime environments, capacity planning, and SLA/SLO management.
  • Experience designing and scaling Proof of Concepts (POCs) into enterprise-grade solutions.
  • Familiarity with Agile methodologies and DevOps practices.
  • Excellent troubleshooting, analytical, and problem-solving skills.

Preferred Qualifications

  • Experience with Artificial Intelligence (AI) technologies.
  • Hands-on knowledge of Generative AI platforms and tools such as:
    • Google Cloud AI (Vertex AI/GHCP)
    • Claude
    • AWS Bedrock
  • Experience implementing AI-enabled operational automation.
  • AWS Professional or Specialty Certifications are a plus.

Technical Skills

  • Java
  • Spring Boot
  • Microservices
  • AWS Cloud
  • Jenkins
  • CloudBees
  • Splunk
  • Datadog
  • Dynatrace
  • New Relic
  • PCF (Pivotal Cloud Foundry)
  • App Pilot
  • Presto
  • CI/CD
  • DevOps
  • Site Reliability Engineering (SRE)
  • Capacity Planning
  • SLA/SLO
  • AI
  • Generative AI
  • AWS Bedrock
  • Claude
  • Google Cloud AI
  • Agile

This is a remote position.

Before You Apply
️
πŸ‡ΊπŸ‡Έ Be aware of the location restriction for this remote position: USA Only
β€Ό Beware of scams! When applying for jobs, you should NEVER have to pay anything. Learn more.
Site Reliability Engineering SRE Leader with AI Experience @ZENITH INFOTEK LLC
Devops
Salary unspecified
Remote Location
πŸ‡ΊπŸ‡Έ USA Only
Employment Type contract
Posted 2mths ago
Apply for this position
Did not apply βœ“
Applied βœ“
Sent Follow-Up βœ“
Interview Scheduled βœ“
Interview Completed βœ“
Offer Accepted βœ“
Offer Declined βœ“
Application Denied βœ“
Unlock 125,000+ Remote Jobs
️
πŸ‡ΊπŸ‡Έ Be aware of the location restriction for this remote position: USA Only
β€Ό Beware of scams! When applying for jobs, you should NEVER have to pay anything. Learn more.
Apply for this position
Did not apply βœ“
Applied βœ“
Sent Follow-Up βœ“
Interview Scheduled βœ“
Interview Completed βœ“
Offer Accepted βœ“
Offer Declined βœ“
Application Denied βœ“
Unlock 125,000+ Remote Jobs
Γ—
Apply to the best remote jobs
before everyone else

Access 125,000+ vetted remote jobs and get daily alerts.

4.9 β˜…β˜…β˜…β˜…β˜… from 500+ reviews

⚑ 127,104+ remote jobs, refreshed hourly

πŸ”” Real-time alerts: Apply first, direct to employer

πŸ›‘οΈ Vetted companies, no scams, true remote only

Unlock All Jobs Now

Maybe later