In short
Document understanding datasets need character level text, layout regions, table structure, field to value relationships and coverage of poor scans, handwriting and multilingual forms.

Character recognition is largely solved for clean print. The remaining difficulty is structure: which text is a heading, which cell belongs to which column, which value answers which field.
What to annotate
- Text with reading order preserved
- Layout regions such as headers, tables, stamps and signatures
- Table structure including merged and spanning cells
- Field and value pairs for the data your workflow must extract
Collect the difficult scans
Real documents are creased, photographed at an angle, partly handwritten and sometimes printed on a form from twenty years ago. A dataset of clean scans creates a workflow that fails on the first day of use.
Multilingual forms
Many official documents mix two languages on the same page. Annotators need clear rules for language tagging so downstream extraction does not silently drop one of them.
Practical checklist
- Define the acceptance rule before any volume starts
- Review a small pilot before committing the full budget
- Track errors by category, language and reviewer
- Keep consent, source notes and version history with the files
Before you ask for a quote
A clear brief saves days. Share a sample file, target language or region, expected volume, deadline, quality threshold and any privacy restrictions. A supplier can then price the work on real effort rather than assumptions.
- Which languages, markets or user groups must be represented?
- What format does the final file need to arrive in?
- Who will approve ambiguous cases during the pilot?
- ocr
- documents
- extraction
Need this done rather than read about it?
We run collection, annotation, transcription and localization projects for teams who would rather spend their time on the product.
Start a conversation