Senior Site Reliability Engineer @EPAM Systems
Devops
Salary unspecified
Remote Location
Employment Type full-time
Posted 2d ago

[Hiring] Senior Site Reliability Engineer @EPAM Systems

2d ago - EPAM Systems is hiring a remote Senior Site Reliability Engineer. πŸ’Έ Salary: unspecified πŸ“Location: Argentina

Role Description

We are looking for a hands-on Senior Site Reliability Engineer to help maintain, enhance, and support a Java services ecosystem in close collaboration with an SRE peer and a backend engineering team. You will strengthen reliability, observability, and operational readiness while participating in on-call support.

  • Provide on-call support for Java backend identity services during business hours
  • Troubleshoot complex production issues using logs and telemetry and drive root-cause resolution
  • Prepare and deploy patches to address issues in cloud infrastructure
  • Improve service reliability by implementing practical changes that reduce errors and instability
  • Build and refine metrics and dashboards to surface platform health and service behavior
  • Monitor SLOs and propose code changes that improve SLO attainment as issues arise
  • Create and improve runbooks to standardize operational response and reduce time to recovery
  • Communicate incidents and operational risks clearly in writing during live response
  • Collaborate closely with engineers to align operational practices with service ownership

Qualifications

  • 3+ years of Site Reliability Engineering or DevOps experience supporting distributed systems
  • Strong on-call support experience for production services and incident response during business hours
  • Proven experience with Amazon Web Services in production environments
  • Hands-on experience with Amazon DynamoDB and Amazon ElastiCache
  • Strong Git skills for collaborating on operational and reliability code changes
  • Solid Gradle knowledge for building and maintaining Java-based services
  • Strong troubleshooting skills using logs and telemetry to identify root causes
  • Clear written communication skills for documenting and reporting operational issues during incidents
  • Proactive learning mindset to absorb complex information quickly and apply it under pressure
  • Upper-Intermediate English proficiency (B2)

Requirements

  • Kubernetes (Nice to have)
  • Terraform (Nice to have)
  • Grafana (Nice to have)
  • New Relic (Nice to have)
  • Apache Kafka (Nice to have)
Before You Apply
️
remote Be aware of the location restriction for this remote position: Argentina
β€Ό Beware of scams! When applying for jobs, you should NEVER have to pay anything. Learn more.
Senior Site Reliability Engineer @EPAM Systems
Devops
Salary unspecified
Remote Location
Employment Type full-time
Posted 2d ago
Apply for this position
Did not apply βœ“
Applied βœ“
Sent Follow-Up βœ“
Interview Scheduled βœ“
Interview Completed βœ“
Offer Accepted βœ“
Offer Declined βœ“
Application Denied βœ“
Unlock 125,000+ Remote Jobs
️
remote Be aware of the location restriction for this remote position: Argentina
β€Ό Beware of scams! When applying for jobs, you should NEVER have to pay anything. Learn more.
Apply for this position
Did not apply βœ“
Applied βœ“
Sent Follow-Up βœ“
Interview Scheduled βœ“
Interview Completed βœ“
Offer Accepted βœ“
Offer Declined βœ“
Application Denied βœ“
Unlock 125,000+ Remote Jobs
Γ—
Apply to the best remote jobs
before everyone else

Access 125,000+ vetted remote jobs and get daily alerts.

4.9 β˜…β˜…β˜…β˜…β˜… from 500+ reviews

⚑ 129,624+ remote jobs, refreshed hourly

πŸ”” Real-time alerts: Apply first, direct to employer

πŸ›‘οΈ Vetted companies, no scams, true remote only

Unlock All Jobs Now

Maybe later