Data Scientist — LLM Evaluation
full-time
mid
Posted 4 months ago
Before you apply
Build my evidence-backed draft — freePaste your relevant resume section or 2–4 true bullets. No generic form is sent; any later submission requires your approval of the exact draft and role.
About this role
Design and implement evaluation frameworks for large language models. Build benchmarks, run experiments, and measure model quality across dimensions.
Your work determines which models ship and which don't.
Requirements
Strong statistics background. Experience with LLM evaluation or NLP benchmarking. Python required. Experience with statistical testing.
Similar Jobs
Related searches:
Get jobs like this delivered weekly
Free AI jobs newsletter. No spam.