In short
Inter annotator agreement measures how consistently reviewers apply the same labels. Low agreement usually indicates unclear guidelines or an unsuitable label set rather than poor annotator performance.

If two trained people label the same item differently, the instructions are ambiguous. The label set may also be asking for a distinction that does not reliably exist.
How to run the check
- Give the same one hundred items to at least three annotators
- Use a chance corrected measure rather than raw percentage match
- Compute agreement per label, since one category usually causes most of the disagreement
- Repeat monthly, because agreement drifts
Reading the result
A single weak category means the guideline for it needs examples. Broad weakness across categories means the label set is too fine grained for reliable human judgement and should be simplified.
Keep the disagreements
Items reviewers disagree on are the best practice material for new annotators and the clearest guideline examples. Store them rather than resolving and discarding them.
Practical checklist
- Define the acceptance rule before any volume starts
- Review a small pilot before committing the full budget
- Track errors by category, language and reviewer
- Keep consent, source notes and version history with the files
Before you ask for a quote
A clear brief saves days. Share a sample file, target language or region, expected volume, deadline, quality threshold and any privacy restrictions. A supplier can then price the work on real effort rather than assumptions.
- Which languages, markets or user groups must be represented?
- What format does the final file need to arrive in?
- Who will approve ambiguous cases during the pilot?
- quality
- metrics
- guidelines
Need this done rather than read about it?
We run collection, annotation, transcription and localization projects for teams who would rather spend their time on the product.
Start a conversation