Senior AI QA Test Automation Engineer @XM Careers
Artificial Intelligence
Salary unspecified
Remote Location
Employment Type full-time
Posted 1wk ago

[Hiring] Senior AI QA Test Automation Engineer @XM Careers

1wk ago - XM Careers is hiring a remote Senior AI QA Test Automation Engineer. πŸ’Έ Salary: unspecified πŸ“Location: Greece

Role Description

We're seeking a Senior AI QA Test Automation Engineer to take a leading technical role in AI quality and test automation. You will work primarily with our AWS Bedrock AgentCore-based CX Agent Suite, a multi-agent customer-support system covering intent routing, FAQs, deposit-status queries, and human handoff. You will design and build the evaluation and QA tooling required to validate AI agents reliably, while also contributing to technical standards and best practices across QA.

This is a Senior Individual Contributor (IC) role with a strong technical and strategic focus. You'll work closely with QA leadership, Data Science, and Engineering to architect scalable AI evaluation systems, establish technical standards, and drive the evolution of our AI evaluation framework. Our current evaluation harness is based on DeepEval and is evolving towards a broader agentic evaluation architecture. You will have the opportunity to evaluate and prototype emerging agent-orchestration approaches, including Strands Agents, and contribute to architectural decisions around their adoption.

  • Act as the primary technical enabler for QA, building scalable AI/ML frameworks, libraries, and tooling for the CX Agent Suite and beyond
  • Design and build AI evaluation pipelines that assess the CX Agent Suite's multi-agent responses (Intent Detector, FAQ Agent, Missing-Deposits Agent) for accuracy, relevance, tone, hallucination rate, safety/guardrail compliance, and task completion β€” extending or replacing the current DeepEval-based evaluation harness
  • Evaluate and prototype Strands Agents (or comparable agent-orchestration frameworks) for building self-evolving, autonomous QA agents, and drive the go/no-go decision on adoption
  • Collaborate with QA, Data Science, and Engineering to integrate AI-driven testing into the existing GitLab CI/CD pipeline (dev β†’ test β†’ staging β†’ prod), alongside Terraform-provisioned, EKS/AgentCore-hosted services
  • Build resilient and adaptive automation that can detect and respond to changes in agent behavior, Bedrock Guardrails configuration, and routing logic
  • Develop data-driven quality analytics, including root-cause analysis, quality trends, and intelligent test prioritization, leveraging existing observability and evaluation data from OpenTelemetry, CloudWatch, X-Ray traces, and LangFuse offline evaluation runs
  • Lead research into emerging AI testing methodologies and mentor engineers through code reviews, workshops, and architectural guidance

Qualifications

  • BSc/MSc in Computer Science, AI, or a related discipline
  • 6+ years of hands-on experience in QA/Test Automation, with strong experience designing and maintaining automation frameworks
  • 1+ years of experience applying AI/ML in software testing or QA process improvement, with a proven track record of bringing AI/ML solutions into production workflows
  • Strong Python skills, with experience working in Python/AWS-centric technology environments; Java and/or TypeScript is a plus
  • Hands-on experience with LLM evaluation techniques, including LLM-as-a-judge, human-in-the-loop evaluation, RAG, and multi-agent orchestration patterns
  • Practical experience with DeepEval or comparable evaluation frameworks, along with familiarity with agent-orchestration SDKs such as Strands Agents
  • Knowledge of LLM tooling, vector databases, and MLOps pipelines
  • Experience integrating AI tooling into enterprise CI/CD environments, particularly GitLab, and working with containerized cloud-native environments such as Docker and Kubernetes/EKS
  • Strong communicator, able to influence technical decisions and drive engineering standards across teams

Requirements

  • Experience with autonomous QA agents or agentic orchestration frameworks for self-evolving test suites
  • Experience with LLM observability tools such as LangFuse, LangSmith, or Arize, particularly for measuring probabilistic and adversarial robustness
  • Knowledge of AI ethics, fairness, and bias detection for guardrail and model validation
  • Experience with gRPC, WebSockets, and/or HTTP/2
  • Experience with AWS Bedrock, plus familiarity with GCP Vertex AI and/or Azure AI

Benefits

  • Attractive remuneration package
  • Intellectually stimulating work environment
  • Continuous personal development and international training opportunities

Company Description

All applications will be treated with strict confidentiality! We may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses and identifying potential inconsistencies or verification signals in application materials based on available information. These tools assist our recruitment team but do not replace human judgment. Final hiring decisions are ultimately made by humans. If you would like more information about how your data is processed, please contact us.

Before You Apply
️
remote Be aware of the location restriction for this remote position: Greece
β€Ό Beware of scams! When applying for jobs, you should NEVER have to pay anything. Learn more.
Senior AI QA Test Automation Engineer @XM Careers
Artificial Intelligence
Salary unspecified
Remote Location
Employment Type full-time
Posted 1wk ago
Apply for this position
Did not apply βœ“
Applied βœ“
Sent Follow-Up βœ“
Interview Scheduled βœ“
Interview Completed βœ“
Offer Accepted βœ“
Offer Declined βœ“
Application Denied βœ“
Unlock 120,000+ Remote Jobs
️
remote Be aware of the location restriction for this remote position: Greece
β€Ό Beware of scams! When applying for jobs, you should NEVER have to pay anything. Learn more.
Apply for this position
Did not apply βœ“
Applied βœ“
Sent Follow-Up βœ“
Interview Scheduled βœ“
Interview Completed βœ“
Offer Accepted βœ“
Offer Declined βœ“
Application Denied βœ“
Unlock 120,000+ Remote Jobs
Γ—
Apply to the best remote jobs
before everyone else

Access 120,000+ vetted remote jobs and get daily alerts.

4.9 β˜…β˜…β˜…β˜…β˜… from 500+ reviews

⚑ 124,780+ remote jobs, refreshed hourly

πŸ”” Real-time alerts: Apply first

πŸ›‘οΈ Vetted companies, no scams, true remote only

Unlock All Jobs Now

Maybe later