In short
Generated examples work well for rare events, privacy sensitive domains and balancing under represented classes. They fail when used as a full substitute for real data because they only reflect patterns already present elsewhere.

Generated examples have moved from research demos to production workflows. Used well, they close gaps that collection cannot reach. Used as a shortcut, they produce products that look fine in testing and fall over in the field.
Good reasons to generate
- Rare events that are dangerous or impractical to record
- Sensitive domains where real records cannot leave a controlled system
- Class balancing when one category is genuinely scarce
- Early prototyping before a collection budget is approved
The gap that catches teams out
A generator produces variation it already knows about. It will not invent the odd camera angle, the regional accent or the paper form nobody standardised. Products prepared mostly on generated examples inherit that narrowness and score well on generated test sets, which hides the problem.
A workable ratio
Keep real collected data as the base of the set and the whole of the evaluation sample. Add generated examples to fill named gaps, label them clearly in your metadata, and compare performance with and without them before deciding to keep them.
Practical checklist
- Define the acceptance rule before any volume starts
- Review a small pilot before committing the full budget
- Track errors by category, language and reviewer
- Keep consent, source notes and version history with the files
Before you ask for a quote
A clear brief saves days. Share a sample file, target language or region, expected volume, deadline, quality threshold and any privacy restrictions. A supplier can then price the work on real effort rather than assumptions.
- Which languages, markets or user groups must be represented?
- What format does the final file need to arrive in?
- Who will approve ambiguous cases during the pilot?
- generated examples
- privacy
- edge cases
Need this done rather than read about it?
We run collection, annotation, transcription and localization projects for teams who would rather spend their time on the product.
Start a conversation