[Hiring] AI QA & Evaluation Engineer @Elastic
AI QA & Evaluation Engineer @Elastic
Quality Assurance
Salary unspecified
Remote Location
Employment Type full-time
Posted 1wk ago

[Hiring] AI QA & Evaluation Engineer @Elastic

1wk ago - Elastic is hiring a remote AI QA & Evaluation Engineer. πŸ’Έ Salary: unspecified πŸ“Location: Worldwide

Role Description

We are looking for a skilled QA & Evaluation Engineer to join our team. The role blends strategic QA leadership with hands-on technical validation and structured evaluation to safeguard the accuracy, reliability, compliance, and ethical use of AI models. You will partner across IT and Engineering teams to identify, design, implement and run robust testing frameworks and evaluation rubrics for a portfolio of GenAI solutions that will be used across our organization.

What You Will Be Doing

  • Be a primary contributor to our AI strategy, helping validate and test AI infrastructure, custom solutions and third-party SaaS offerings.
  • Test Strategy & Execution:
    • Design and implement comprehensive test strategies for AI/ML systems, including accuracy, bias, robustness, and regression testing.
  • Rubric-Based Evaluation:
    • Design and implement self-contained evaluation tasks, including prompts, supporting files, and detailed grading rubrics to assess AI performance on functional workflows.
  • Automation & CI/CD:
    • Automate validation suites for agentic/multi-agent systems, integration testing, and CI/CD pipelines for ML models.
  • Data Validation:
    • Validate that AI/ML models are consuming accurate, authorized, and properly structured data sources; ensuring data quality across training and inference.
  • Observation & Reporting:
    • Meticulously observe and document AI agent behaviors, producing crisp, precise summaries and reports on model performance and hallucinations.
  • Output Grounding:
    • Validate prompt engineering outputs from a data accuracy standpoint, ensuring responses are grounded in verified data sources.
  • Refinement & Iteration:
    • Iterate and refine evaluation tasks and rubrics based on feedback and team collaboration to ensure robust benchmarking methodologies.
  • Security & Governance:
    • Ensure all AI data sources and structures meet governance, regulatory, and compliance standards, while implementing best practices for security and data privacy.
  • Collaborate with teams from different areas including IT Engineering, IT Operations, Data & Integrations, PMO, CRM, Risk & Compliance, and business technology.
  • Stay current on the latest work in AI and make technical recommendations to the organization.

Qualifications

  • Proficiency in Python, TypeScript, or other programming languages used in AI and test automation.
  • Proven skill in designing or applying rubric-based evaluation, grading against set criteria, or building structured scoring frameworks.
  • Direct experience with LLM evaluation frameworks and benchmarking tools such as LangSmith, Confident AI, etc.
  • Knowledge of the GenAI stack and solutions including Retrieval Augmented Generation (RAG).
  • LLMs: Azure OpenAI, Vertex AI, ChatGPT Enterprise or similar.
  • High attention to detail and ability to notice subtle patterns or inconsistencies (such as data hallucinations or logic errors) that others might miss.
  • Advanced written communication skills, especially for documenting nuanced observations and feedback.
  • Experience with Cloud platforms (Azure, GCP, AWS).
  • Thorough understanding of DevOps/automation/CI/CD tools: GitHub, Terraform.
  • Comprehension with AI ethics, risk management, and data governance.

Benefits

  • Competitive pay based on the work you do here and not your previous salary.
  • Health coverage for you and your family in many locations.
  • Ability to craft your calendar with flexible locations and schedules for many roles.
  • Generous number of vacation days each year.
  • Increase your impact - We match up to $2000 (or local currency equivalent) for financial donations and service.
  • Up to 40 hours each year to use toward volunteer projects you love.
  • Embracing parenthood with a minimum of 16 weeks of parental leave.
Before You Apply
️
worldwide Be aware of the location restriction for this remote position: Worldwide
β€Ό Beware of scams! When applying for jobs, you should NEVER have to pay anything. Learn more.
AI QA & Evaluation Engineer @Elastic
Quality Assurance
Salary unspecified
Remote Location
Employment Type full-time
Posted 1wk ago
Apply for this position
Did not apply βœ“
Applied βœ“
Sent Follow-Up βœ“
Interview Scheduled βœ“
Interview Completed βœ“
Offer Accepted βœ“
Offer Declined βœ“
Application Denied βœ“
Unlock 120,000+ Remote Jobs
️
worldwide Be aware of the location restriction for this remote position: Worldwide
β€Ό Beware of scams! When applying for jobs, you should NEVER have to pay anything. Learn more.
Apply for this position
Did not apply βœ“
Applied βœ“
Sent Follow-Up βœ“
Interview Scheduled βœ“
Interview Completed βœ“
Offer Accepted βœ“
Offer Declined βœ“
Application Denied βœ“
Unlock 120,000+ Remote Jobs