What is data labelling?
A clear guide to classification, annotation, review and the quality controls behind labelled datasets.
4 min readData labelling adds structured meaning to raw information so it can be used for training or evaluation. A label might identify an object, classify intent, score an answer or mark a policy violation.
Examples of labelling work
The interface changes with the data, but each task should connect an input to an explicit decision.
- Text classification and sentiment
- Image regions and object labels
- Transcription and entity extraction
- Response scoring and preference ranking
- Safety and policy review
Where quality breaks down
Ambiguous labels, missing edge cases and inconsistent reviewer interpretation can damage a dataset. Pilot batches and disagreement review should happen before volume production.
Professional labelling
Some decisions require more than annotation experience. Medical, legal, financial and technical content may need contributors who understand the consequences of a wrong label.
Designing a label system
A useful taxonomy is mutually understandable, covers the intended population and gives contributors a path for genuinely ambiguous cases. Begin with the decision the model or analyst must make, then define the smallest label set that supports it. More categories increase training and review cost and may reduce agreement.
Annotation workflow from pilot to delivery
Teams should pilot instructions on a varied sample, compare decisions, revise unclear definitions and only then scale. Production batches need qualification, quality sampling, adjudication and versioned exports so corrected records can be traced without silently changing prior deliveries.
Metrics that reveal real quality
Inter-annotator agreement is useful but incomplete. Pair it with gold accuracy, per-label confusion, abstention rates, reviewer corrections and coverage. Inspect the underlying examples whenever a metric changes; one rare but consequential category may matter more than the overall average.
Privacy, bias and worker context
Remove data that is not needed for the task and restrict access to sensitive material. Test whether labels encode cultural assumptions or systematically treat groups differently. Contributors also need warnings, breaks and support when projects contain distressing material.