Senior AI Reliability Engineer @EPAM Systems
Artificial Intelligence
Salary unspecified
Remote Location
Employment Type full-time
Posted 4wks ago

[Hiring] Senior AI Reliability Engineer @EPAM Systems

4wks ago - EPAM Systems is hiring a remote Senior AI Reliability Engineer. πŸ’Έ Salary: unspecified πŸ“Location: Ukraine

Role Description

EPAM's Operational Intelligence practice is developing a new capability called AI Reliability Engineering (AIRE), which applies SRE principles and cloud-native practices throughout the lifecycle of production AI/ML and LLM systems. This role exists to close the gap in visibility into token latency, cost per request, semantic drift, or hallucinations as clients transition their GenAI applications and agentic systems from pilot to production.

Your work will involve:

  • Instrumenting, monitoring, and hardening production AI systems.
  • Establishing AI-native service level objectives.
  • Creating accelerators and reference implementations for reuse across accounts.

Note that this is an engineering position rather than an L1/L2 support role, and it does not require a 24/7 on-call rotation.

Qualifications

  • 4+ years of experience in SRE, DevOps, platform, or observability engineering, with hands-on exposure to production AI/ML or LLM workloads.
  • Strong grasp of SRE fundamentals, including Golden Signals, SLI/SLO definition, error budgets and burn rate, incident lifecycle, and ITIL basics.
  • Strong Python skills for building instrumentation, automation, and evaluation tools.
  • Hands-on experience with OpenTelemetry and at least one APM/observability platform such as New Relic, Datadog, Grafana LGTM stack, Splunk, or Elastic.
  • Production experience with at least one cloud platform (Azure preferred, AWS or GCP also acceptable) and Kubernetes.
  • Solid understanding of LLM application architecture, including prompts, embeddings and vector stores, RAG, and agent orchestration tools like LangChain/LangGraph or similar.
  • Awareness of MLOps concepts, including model lifecycle (training vs. inference), model endpoints, containerization, and deployment/rollback patterns.
  • Experience with Infrastructure as Code using Terraform, along with CI/CD tools such as Azure DevOps, GitLab CI, or GitHub Actions.
  • B2+ English proficiency, since the role involves direct client interaction and requires clear technical communication in both writing and speech.

Requirements

  • Add AI telemetry to production LLM, RAG, and agentic applications using OpenTelemetry and APM-native AI monitoring tools.
  • Deploy distributed tracing across multi-model chains, agent workflows, and retrieval-augmented generation pipelines to identify systemic latency and failure points.
  • Establish and track AI-native SLIs and SLOs, including time to first token (TTFT), throughput, error and refusal rates, cost per request, semantic drift, hallucination boundaries, and contextual accuracy.
  • Implement structured semantic logging and prompt/response monitoring to support quality analysis.
  • Create and maintain evaluation loops for output quality and safety using golden sets, LLM-as-a-judge methods, and Ragas/DeepEval-style frameworks, integrating them into CI/CD and runtime environments.
  • Monitor token-based cloud spend, model API rate limits, and quota usage while driving AI cost optimization efforts.
  • Set up AI gateways to manage API load balancing, failover, and fallback models across multiple LLM providers.
  • Build guardrails covering prompt injection and jailbreak filtering, output compliance, and bias and safety constraints.
  • Develop detection, triage, restoration, and problem management workflows for AI incidents, incorporating autonomous AI agents into root cause analysis to parse logs, generate hypotheses, and correlate state changes.
  • Enable rollback, canary, and fail-safe patterns for model, prompt, and configuration releases, while maintaining reproducibility through versioning of data, code, prompts, and models.
  • Develop practice accelerators, reference architectures, and internal training materials, and contribute to presales activities and client assessments.

Benefits

  • A brand-new discipline within EPAM where you shape the approach rather than follow an existing playbook.
  • Focus on engineering work without 24/7 on-call responsibilities.
  • Sponsored certification and training programs (Anthropic/Claude, Databricks, AI & Data Observability learning paths).
  • Exposure to multiple clients and a direct route into presales and solution engineering.
Before You Apply
️
remote Be aware of the location restriction for this remote position: Ukraine
β€Ό Beware of scams! When applying for jobs, you should NEVER have to pay anything. Learn more.
Senior AI Reliability Engineer @EPAM Systems
Artificial Intelligence
Salary unspecified
Remote Location
Employment Type full-time
Posted 4wks ago
Apply for this position
Did not apply βœ“
Applied βœ“
Sent Follow-Up βœ“
Interview Scheduled βœ“
Interview Completed βœ“
Offer Accepted βœ“
Offer Declined βœ“
Application Denied βœ“
Unlock 125,000+ Remote Jobs
️
remote Be aware of the location restriction for this remote position: Ukraine
β€Ό Beware of scams! When applying for jobs, you should NEVER have to pay anything. Learn more.
Apply for this position
Did not apply βœ“
Applied βœ“
Sent Follow-Up βœ“
Interview Scheduled βœ“
Interview Completed βœ“
Offer Accepted βœ“
Offer Declined βœ“
Application Denied βœ“
Unlock 125,000+ Remote Jobs
Γ—
Apply to the best remote jobs
before everyone else

Access 125,000+ vetted remote jobs and get daily alerts.

4.9 β˜…β˜…β˜…β˜…β˜… from 500+ reviews

⚑ 126,845+ remote jobs, refreshed hourly

πŸ”” Real-time alerts: Apply first, direct to employer

πŸ›‘οΈ Vetted companies, no scams, true remote only

Unlock All Jobs Now

Maybe later