Build and annotate the datasets a model needs in order to learn.
Next intake: Monday 5 October, with enrolment open until 2 October. The following one starts on 4 January 2027.
Understand each stage of the training data life cycle and who is involved
Annotate text, images and audio methodically with a professional tool
Measure agreement between annotators and resolve a disagreement
Write an annotation guide that a third party can apply without calling you
Clean a dataset: duplicates, missing values, outliers
Anonymise or pseudonymise according to what the intended use allows
Deliver a versioned dataset with its datasheet and terms of reuse
Systematic alternation between short theory sessions and practical work, with most of the time spent on practice
Progression from simple to complex: each module builds on what was learnt in the previous one
Realistic role-plays drawn from Cameroonian business cases
Considered use of generative AI as a working tool, in line with the common core
Regular formative assessments and a final integrative project drawing on all four modules
Formative assessment at the end of each module (graded role-play and practical exercises)
Practical work assessed against competency sheets
Final integrative project, presented and defended before the trainer
One graded assignment per module in the online course space
Label Studio or CVAT — text, image and audio annotation
Spreadsheet and basic Python for cleaning
Audacity for audio segmentation
Git and object storage for dataset versioning
Public datasets and anonymised company datasets
Sessions open every quarter. Apply now or request the detailed brochure — our advisers will get back to you.