AI Evaluation & Benchmarking Engineer @Intergral
Artificial Intelligence
Salary gbp 60,000 - 75..
Remote Location
remote UK
Employment Type full-time
Posted 2wks ago

[Hiring] AI Evaluation & Benchmarking Engineer @Intergral

2wks ago - Intergral is hiring a remote AI Evaluation & Benchmarking Engineer. πŸ’Έ Salary: gbp 60,000 - 75,000 per year πŸ“Location: UK

Role Description

We're looking for an AI Evaluation & Benchmarking Engineer to determine how we measure the quality of OpsPilot, and to build the systems that do it.

More important than any individual technology is how you approach measurement. We're looking for someone who:

  • Questions whether a system is actually achieving the outcome it was designed for.
  • Works out how to measure that objectively.
  • Builds what's needed to keep measuring it as the product changes.

You'll have real autonomy over the technical approach. We set the goals and check in regularly, but the expertise on how to get there is yours. This is a hands-on engineering role with no reports and no QA silo. Our engineering structure is flat: you'll report to the Director of Engineering and work alongside the other engineers as a peer.

This isn't a traditional QA role. You won't be manually testing tickets, acting as a release gatekeeper or simply checking whether features technically work.

What You'll Do

  • Evaluate the agent and the product it runs on.
  • On the agentic side, you'll:
    • Build automated evaluations for our agentic AI capabilities.
    • Create realistic synthetic scenarios, datasets and workloads with meaningful ground truth and evaluation criteria.
    • Measure task success, diagnostic accuracy, evidence quality, reliability, consistency, latency and cost.
    • Account for the non-deterministic nature of AI systems through repeated runs, variance analysis and determining whether changes are meaningful rather than noise.
    • Benchmark models, prompts, tools, retrieval strategies and agent workflows against repeatable baselines.
    • Use relevant industry benchmarks and standards, including established SRE practices and emerging AI-agent, AIOps and incident-response benchmarks, and build our own where existing approaches don't represent real operational work.
  • Across the wider product, you'll:
    • Extend evaluation across important customer journeys, APIs, backend services and UI.
    • Use OpenTelemetry, including its Semantic Conventions, to make benchmark environments representative and portable.
    • Use product telemetry and correlate benchmark results with metrics, logs and traces to understand why failures occur.
    • Build end-to-end measurements focused on customer outcomes rather than isolated components.
    • When we change a model, prompt, tool or agent workflow, we want to know what became better, what became worse and why β€” including the impact on quality, reliability, latency and cost.
  • Turn what we learn into continuous improvement:
    • Turn failures and real-world problems into new evaluation scenarios.
    • Identify recurring failure patterns and capability gaps.
    • Test potential improvements against established baselines, holdouts and unseen scenarios.
    • Detect regressions, benchmark overfitting and improvements that don't generalise.
    • Longer term, this evaluation system becomes the harness for controlled self-improvement: identifying weaknesses, testing potential changes and objectively determining whether they should be retained.
  • Work with the rest of engineering:
    • Work directly with AI, software, platform and SRE engineers to investigate findings and improve the product.
    • Make evaluation failures clear, reproducible and actionable.
    • Make straightforward fixes yourself where that's the most efficient approach.
    • Build tooling that makes evaluations easy for other engineers to create, run and understand.
    • Use AI-assisted engineering where it improves the speed or quality of your work.

Qualifications

  • Ability to take an ambiguous technical problem, develop an approach and deliver a working system independently.
  • 3+ years of relevant technical experience (background as an AI engineer, software engineer, SRE, platform engineer, performance engineer, SDET or similar).
  • Practical experience working with LLMs, AI agents or AI evaluation.
  • Strong software engineering skills, particularly in Python or a similar language.
  • Experience building automated evaluation, benchmarking, testing or experimentation infrastructure.
  • Experience creating synthetic workloads, datasets or evaluation scenarios.
  • Understanding of non-deterministic evaluation, including repeated measurement, variance and distinguishing meaningful changes from noise.
  • Ability to turn complex system behaviour into measurable criteria.
  • Comfort working across APIs, distributed systems and multiple layers of a software product.

Desirable

  • Agentic AI evaluation, tool use and multi-step workflows.
  • LLM evaluation frameworks and model-based evaluation techniques.
  • Automated experimentation or self-improving systems.
  • Dataset, ground-truth and holdout evaluation design.
  • Statistical experimentation and performance benchmarking.
  • OpenTelemetry, metrics, logs and distributed tracing.
  • SRE, incident response or observability.
  • Production SaaS and distributed systems.

What We Offer

  • A small company with under ten people in engineering and a flat structure.
  • Fully remote within the UK.
  • Flexible working hours.
  • 25 days holiday plus bank holidays.
  • Real autonomy over your technical approach and how you deliver the role.

Our Interview Process

Straightforward: usually two or three conversations, with no technical coding tests.

If this sounds like the kind of challenge you're looking for, we'd love to hear from you.

Before You Apply
️
remote Be aware of the location restriction for this remote position: UK
β€Ό Beware of scams! When applying for jobs, you should NEVER have to pay anything. Learn more.
AI Evaluation & Benchmarking Engineer @Intergral
Artificial Intelligence
Salary gbp 60,000 - 75..
Remote Location
remote UK
Employment Type full-time
Posted 2wks ago
Apply for this position
Did not apply βœ“
Applied βœ“
Sent Follow-Up βœ“
Interview Scheduled βœ“
Interview Completed βœ“
Offer Accepted βœ“
Offer Declined βœ“
Application Denied βœ“
Unlock 125,000+ Remote Jobs
️
remote Be aware of the location restriction for this remote position: UK
β€Ό Beware of scams! When applying for jobs, you should NEVER have to pay anything. Learn more.
Apply for this position
Did not apply βœ“
Applied βœ“
Sent Follow-Up βœ“
Interview Scheduled βœ“
Interview Completed βœ“
Offer Accepted βœ“
Offer Declined βœ“
Application Denied βœ“
Unlock 125,000+ Remote Jobs
Γ—
Apply to the best remote jobs
before everyone else

Access 125,000+ vetted remote jobs and get daily alerts.

4.9 β˜…β˜…β˜…β˜…β˜… from 500+ reviews

⚑ 126,955+ remote jobs, refreshed hourly

πŸ”” Real-time alerts: Apply first, direct to employer

πŸ›‘οΈ Vetted companies, no scams, true remote only

Unlock All Jobs Now

Maybe later