Evaluations Team Lead @Fundamental
Artificial Intelligence
Salary unspecified
Remote Location
Employment Type full-time
Posted YDay

[Hiring] Evaluations Team Lead @Fundamental

YDay - Fundamental is hiring a remote Evaluations Team Lead. πŸ’Έ Salary: unspecified πŸ“Location: USA, Israel

Role Description

NEXUS is already in production, and we are working to substantially improve its predictive quality, latency, and cost efficiency. You will lead the team that owns how NEXUS is measured, building the shared evaluation platform that engineering, research, and Applied AI all rely on:

  • Consistent benchmarks
  • Curated datasets
  • Standards for metrics, data splits, leakage prevention, and benchmark contamination that hold up under scrutiny

Research needs to trust results before committing compute to it, engineering needs regressions caught before they reach production, and when a customer's own data science team benchmarks NEXUS against their own models, your platform is what Applied AI relies upon. You will benchmark NEXUS against competing approaches, using fair tuning budgets, data access, and latency measurement protocols, and turn what you find into research priorities and release recommendations.

This is a player-coach role: you will hire and manage a small team while staying hands-on with the code, the experiment design, and the methodology yourself. There is no evaluation function to inherit here - what you build becomes the standard the rest of the company measures NEXUS against.

Qualifications

  • Experience owning evaluation for tabular ML systems used in production or consequential customer decisions.
  • Strong statistical judgment: choosing metrics and validation schemes, estimating uncertainty, comparing models across datasets, and accounting for repeated experimentation.
  • Practical experience finding leakage in preprocessing, feature construction, joins, temporal dependencies, and related entities across splits.
  • Strong Python and SQL skills, familiarity with scikit-learn and gradient-boosted trees, and experience building reliable ML tooling or platforms used by other teams.
  • Experience designing fair model comparisons, including hyperparameter search, resource budgets, and end-to-end latency measurement.
  • Prior people management experience, including hiring, technical coaching, and performance feedback, while remaining technically involved.
  • Clear written and spoken communication with researchers, engineers, and customer data scientists, including the willingness to challenge claims the evidence does not support.

Requirements

  • Build a shared evaluation platform for engineering, research, and Applied AI.
  • Continuously add and maintain models and curated datasets so teams can run benchmarks and investigate results independently.
  • Define evaluation standards for metrics, data splits, leakage prevention, calibration, uncertainty, and benchmark contamination.
  • Build reproducible pipelines with versioned inputs and artifacts, and integrate regression checks into research and release workflows.
  • Benchmark NEXUS against competing approaches using fair tuning budgets, data access, compute, and latency measurement protocols.
  • Support Applied AI’s customer POC evaluations with tooling, methodological guidance, and analysis.
  • Measure predictive quality, latency, and cost across deployment configurations, task types, and dataset characteristics.
  • Turn findings into research priorities, release recommendations, and evidence-backed customer improvement plans.
  • Hire and develop the team, set priorities, and stay hands-on with code and experimental design.

Benefits

  • Competitive compensation with salary and equity
  • Comprehensive health coverage for you and your dependents
  • Paid parental leave for all new parents, inclusive of adoptive and surrogate journeys
  • Relocation support for employees moving to join the team in one of our office locations
  • A mission-driven, low-ego culture that values diversity of thought, ownership, and bias toward action
Before You Apply
️
remote Be aware of the location restriction for this remote position: USA, Israel
β€Ό Beware of scams! When applying for jobs, you should NEVER have to pay anything. Learn more.
Evaluations Team Lead @Fundamental
Artificial Intelligence
Salary unspecified
Remote Location
Employment Type full-time
Posted YDay
Apply for this position
Did not apply βœ“
Applied βœ“
Sent Follow-Up βœ“
Interview Scheduled βœ“
Interview Completed βœ“
Offer Accepted βœ“
Offer Declined βœ“
Application Denied βœ“
Unlock 130,000+ Remote Jobs
️
remote Be aware of the location restriction for this remote position: USA, Israel
β€Ό Beware of scams! When applying for jobs, you should NEVER have to pay anything. Learn more.
Apply for this position
Did not apply βœ“
Applied βœ“
Sent Follow-Up βœ“
Interview Scheduled βœ“
Interview Completed βœ“
Offer Accepted βœ“
Offer Declined βœ“
Application Denied βœ“
Unlock 130,000+ Remote Jobs
Γ—
Apply to the best remote jobs
before everyone else

Access 130,000+ vetted remote jobs and get daily alerts.

4.9 β˜…β˜…β˜…β˜…β˜… from 500+ reviews

⚑ 131,011+ remote jobs, refreshed hourly

πŸ”” Real-time alerts: Apply first, direct to employer

πŸ›‘οΈ Vetted companies, no scams, true remote only

Unlock All Jobs Now

Maybe later