|
Salary
unspecified
|
Remote
Location
|
|
Employment Type
contract
|
Posted
2wks ago
|
2wks ago - LILT (Production) is hiring a remote AI Benchmark Engineer. πΈ Salary: unspecified πLocation: Belgium
Role Description
We are building a rigorous, verifiable evaluation suite of Terminal-Bench tasks designed to test the limits of large language models on multilingual software challenges. Our goal is to measure multilingual robustness across prompt language effects, non-English data processing, and complex locale/encoding edge cases in terminal workflows.
We are seeking experienced native-speaking software engineers to design, build, and validate these benchmarks. You will create high-signal, high-quality tasks that genuinely test a model's ability to handle multilingual environments without relying on English translation crutches.
Note this is a remote, freelance opportunity.
What Youβll Deliver
Qualifications
Benefits
What to Consider Before Applying
How to join our expert community
|
|
Be aware of the location restriction for this remote position: Belgium |
| βΌ | Beware of scams! When applying for jobs, you should NEVER have to pay anything. Learn more. | οΈ
|
Salary
unspecified
|
Remote
Location
|
|
Employment Type
contract
|
Posted
2wks ago
|
|
|
Be aware of the location restriction for this remote position: Belgium |
| βΌ | Beware of scams! When applying for jobs, you should NEVER have to pay anything. Learn more. | οΈ
Access 130,000+ vetted remote jobs and get daily alerts.
β‘ 130,438+ remote jobs, refreshed hourly
π Real-time alerts: Apply first
π‘οΈ Vetted companies, no scams, true remote only