Objectives
- Understand benchmarks (MMLU and others) and their limits
- Build an evaluation set for your business
- Compare models on quality, cost and latency
- Evaluate continuously in production
Curriculum
- Benchmark landscape3 h
- MMLU: method and reading scores3 h
- Building your own evaluation set4 h
- Comparing quality, cost, latency2 h
- Continuous evaluation2 h
Audience & prerequisites
Audience: Data scientists, AI engineers, CIOs and technical leads who select models.
Prerequisites: Python, machine learning basics.
Tools used
PythonEvaluation setsLLM APIs
Assessment and certification
The programme ends with a practical project assessed by a panel of practitioners. Passing it earns the Africa Data Entry certificate « MMLU & Language Model Evaluation ».
Enrolment
An advisor will contact you within 48 hours to confirm your enrolment and funding.
