|
Salary
unspecified
|
Remote
Location
|
|
Employment Type
temporary
|
Posted
YDay
|
YDay - Whisk is hiring a remote AI Quality & Evaluation Lead. πΈ Salary: unspecified πLocation: Poland
Role Description
We're building an AI coaching experience inside Samsung Health. Every piece of coaching a user sees is generated, not written. Our platform partner owns the evaluation runtime and functional evals (did the pipeline execute, did the tool call succeed). Nobody owns the layer above that: is the output actually good β right framing, right tone, right structure, does it deliver the coaching logic, does it feel like it's talking to me.
This engagement exists to build that layer: author the failure taxonomy, codify it into a binary rubric, validate an LLM judge against your own human grades, and hand the whole thing to an internal owner. We have a v0.1 rubric and golden scenarios from our product lead, and a functional rubric owned by engineering. We don't have the experience rubric β that's what this engagement produces.
Deliverables by Phase (Overall project ~13 weeks)
What This Role Is Not
What We're Looking For
|
|
Be aware of the location restriction for this remote position: Poland |
| βΌ | Beware of scams! When applying for jobs, you should NEVER have to pay anything. Learn more. | οΈ
|
Salary
unspecified
|
Remote
Location
|
|
Employment Type
temporary
|
Posted
YDay
|
|
|
Be aware of the location restriction for this remote position: Poland |
| βΌ | Beware of scams! When applying for jobs, you should NEVER have to pay anything. Learn more. | οΈ
Access 125,000+ vetted remote jobs and get daily alerts.
β‘ 126,674+ remote jobs, refreshed hourly
π Real-time alerts: Apply first, direct to employer
π‘οΈ Vetted companies, no scams, true remote only