AI Quality & Evaluation Lead @Whisk
Artificial Intelligence
Salary unspecified
Remote Location
Employment Type temporary
Posted YDay

[Hiring] AI Quality & Evaluation Lead @Whisk

YDay - Whisk is hiring a remote AI Quality & Evaluation Lead. πŸ’Έ Salary: unspecified πŸ“Location: Poland

Role Description

We're building an AI coaching experience inside Samsung Health. Every piece of coaching a user sees is generated, not written. Our platform partner owns the evaluation runtime and functional evals (did the pipeline execute, did the tool call succeed). Nobody owns the layer above that: is the output actually good β€” right framing, right tone, right structure, does it deliver the coaching logic, does it feel like it's talking to me.

This engagement exists to build that layer: author the failure taxonomy, codify it into a binary rubric, validate an LLM judge against your own human grades, and hand the whole thing to an internal owner. We have a v0.1 rubric and golden scenarios from our product lead, and a functional rubric owned by engineering. We don't have the experience rubric β€” that's what this engagement produces.

Deliverables by Phase (Overall project ~13 weeks)

  • Failure taxonomy and golden set (within 4 weeks from the contract start date)
    • At least 150 traces hand-graded with open-coded notes.
    • At least 40 of them come from sparse-data synthetic profiles.
    • The taxonomy has 5–10 categories, each with a frequency count, and the top three failures are shown with their share of all failures.
    • Saturation is evidenced: the last 20 traces produced no new category.
    • At least 50 traces are frozen as a holdout.
    • Head of Product signs off the taxonomy.
  • Rubric v1 and calibration (within 4 weeks from phase 1 completion)
    • One binary pass/fail criterion per taxonomy category, each with a definition, a pass example and a fail example.
    • Written annotation guidelines.
    • At least 30 traces are double-coded with a second grader, and every disagreement is logged and resolved in writing.
    • A guideline revision log records each change and its reason.
    • Head of Product signs off the rubric.
  • Validated judge and weekly readout (within 4 weeks from phase 2 completion)
    • Judge prompts for every criterion.
    • A validation report giving true positive rate and true negative rate per criterion on both the dev set and the holdout.
    • A weekly readout on helpfulness, relevance and tone, run at least twice: once by the consultant, and once by the internal owner with the consultant shadowing.
    • A documented revalidation routine, triggered by every model or prompt change and run quarterly regardless.
  • Playbook and handover test (within 1 week from phase 3 completion)
    • A written playbook covering the whole method: grading, updating the taxonomy, revising the rubric, revalidating the judge and running the readout.
    • The handover test passes when the internal owner reruns the judge validation unaided, gets within Β±5 points, and runs one weekly readout with no consultant involvement.

What This Role Is Not

  • Not building eval infrastructure, harnesses, or dashboards (owned elsewhere)
  • Not functional/module-level evals or accuracy metrics
  • Not clinical safety validation
  • Not high-volume annotation β€” grading a few hundred traces to build the taxonomy; volume work goes to vendors/the judge once the rubric exists

What We're Looking For

  • You've run this loop end-to-end at least once on conversational or generated-text output: error analysis on real traces β†’ a failure taxonomy you built yourself β†’ binary criteria and annotation guidelines β†’ an LLM judge validated against your own labels.
  • Background: conversation design, AI product quality, model behaviour/policy, content design, human data operations, UX research with strong qualitative coding, or RLHF. Applied linguistics and product management backgrounds also fit.
  • You'll be strong in this role if you:
    • Can describe a failure mode you personally discovered by reading output β€” one nobody had named before you.
    • Are comfortable being the arbiter and documenting what you overruled and why.
    • Understand the rubric is discovered through grading, not written upfront.
    • Can explain why a judge that agrees with you 94% of the time might still be useless.
    • Write annotation guidelines that don't need interpreting.
    • Are comfortable working in notebooks and spreadsheets (no production code required).
  • Not required: domain expertise in nutrition, weight management, or behaviour change β€” we have that in-house.
Before You Apply
️
remote Be aware of the location restriction for this remote position: Poland
β€Ό Beware of scams! When applying for jobs, you should NEVER have to pay anything. Learn more.
AI Quality & Evaluation Lead @Whisk
Artificial Intelligence
Salary unspecified
Remote Location
Employment Type temporary
Posted YDay
Apply for this position
Did not apply βœ“
Applied βœ“
Sent Follow-Up βœ“
Interview Scheduled βœ“
Interview Completed βœ“
Offer Accepted βœ“
Offer Declined βœ“
Application Denied βœ“
Unlock 125,000+ Remote Jobs
️
remote Be aware of the location restriction for this remote position: Poland
β€Ό Beware of scams! When applying for jobs, you should NEVER have to pay anything. Learn more.
Apply for this position
Did not apply βœ“
Applied βœ“
Sent Follow-Up βœ“
Interview Scheduled βœ“
Interview Completed βœ“
Offer Accepted βœ“
Offer Declined βœ“
Application Denied βœ“
Unlock 125,000+ Remote Jobs
Γ—
Apply to the best remote jobs
before everyone else

Access 125,000+ vetted remote jobs and get daily alerts.

4.9 β˜…β˜…β˜…β˜…β˜… from 500+ reviews

⚑ 126,674+ remote jobs, refreshed hourly

πŸ”” Real-time alerts: Apply first, direct to employer

πŸ›‘οΈ Vetted companies, no scams, true remote only

Unlock All Jobs Now

Maybe later