Agent Evaluation Engineer @EPAM
Artificial Intelligence
Salary unspecified
Remote Location
Employment Type full-time
Posted 1wk ago

[Hiring] Agent Evaluation Engineer @EPAM

1wk ago - EPAM is hiring a remote Agent Evaluation Engineer. πŸ’Έ Salary: unspecified πŸ“Location: Portugal

Role Description

We're looking for an Agent Evaluation Engineer - Build-Time Framework & Deployment Gates to join our team in Portugal in a fully remote working mode. In this role, you will design and maintain an evaluation framework for AI agents, ensuring quality and compliance through automated tests and CI/CD deployment gates. You will develop multi-layer evaluation suites that blend deterministic checks with LLM-powered graders, simulate multi-turn conversations, and define reliability metrics. The position also involves implementing staging validations, shadow-mode traffic analysis, and A/B rollout strategies, with feedback loops from production environments to enhance overall system robustness.

  • Design and implement build-time evaluation frameworks for agentic workflows using LangGraph or comparable orchestration frameworks
  • Create deterministic and LLM-as-judge grading pipelines covering reasoning, trajectory accuracy, and output quality
  • Develop test harnesses for multi-turn conversational simulations and context-retention scoring
  • Define reliability assessment methods including multi-trial metrics (pass@k, pass^k)
  • Implement CI/CD deployment gates that enforce quality thresholds and block releases not meeting standards
  • Integrate staging validation, shadow-mode traffic comparison, and A/B rollout control in deployment pipelines
  • Leverage AWS AgentCore Evaluations for on-demand and online scoring components connected to production feedback
  • Convert production incidents into reusable regression cases for continuous quality improvement
  • Collaborate with engineering and DevOps teams to embed evaluation gates into automated workflows

Qualifications

  • 4+ years of experience in building automated testing or evaluation frameworks for ML, LLM, or agentic systems
  • Proven hands-on experience designing multi-layer evaluation suites with deterministic and LLM-based graders
  • Expertise with CI/CD pipelines and implementing metric-based quality gates for automated deployments
  • Practical knowledge of LangGraph or similar agent orchestration frameworks
  • Strong background in designing simulation-based evaluation strategies and conversation-level tests

Requirements

  • Experience with AWS AgentCore Evaluations API (CreateEvaluation, custom evaluators)
  • Familiarity with shadow-mode, canary, or A/B deployment practices for ML-based platforms
  • Background in transforming production failures into build-time regression tests for agent workflows
Before You Apply
️
remote Be aware of the location restriction for this remote position: Portugal
β€Ό Beware of scams! When applying for jobs, you should NEVER have to pay anything. Learn more.
Agent Evaluation Engineer @EPAM
Artificial Intelligence
Salary unspecified
Remote Location
Employment Type full-time
Posted 1wk ago
Apply for this position
Did not apply βœ“
Applied βœ“
Sent Follow-Up βœ“
Interview Scheduled βœ“
Interview Completed βœ“
Offer Accepted βœ“
Offer Declined βœ“
Application Denied βœ“
Unlock 130,000+ Remote Jobs
️
remote Be aware of the location restriction for this remote position: Portugal
β€Ό Beware of scams! When applying for jobs, you should NEVER have to pay anything. Learn more.
Apply for this position
Did not apply βœ“
Applied βœ“
Sent Follow-Up βœ“
Interview Scheduled βœ“
Interview Completed βœ“
Offer Accepted βœ“
Offer Declined βœ“
Application Denied βœ“
Unlock 130,000+ Remote Jobs
Γ—
Apply to the best remote jobs
before everyone else

Access 130,000+ vetted remote jobs and get daily alerts.

4.9 β˜…β˜…β˜…β˜…β˜… from 500+ reviews

⚑ 131,404+ remote jobs, refreshed hourly

πŸ”” Real-time alerts: Apply first

πŸ›‘οΈ Vetted companies, no scams, true remote only

Unlock All Jobs Now

Maybe later