For supervised fine-tuning and domain judgment, the expert writes the gold answer the model learns from, so the author's qualifications are the ceiling on your data quality. This is a guide to the categories of provider, ranked by what makes gold data actually gold. Full disclosure: we build Pathwize, so read this as an informed but interested guide, with the criteria kept explicit.
How we ranked them
Three things decide the value of expert SFT data. Author qualifications: is the person writing the gold answer genuinely credentialed in the field? Nuance: does the process capture edge cases and reasoning, or flatten everything into a binary label? And provenance: can you prove who authored each example, under what instructions, as Annex IV expects? We ranked the categories on all three.
1. Pathwize, best for verifiable, EU-native expert data
Pathwize sources gold answers and expert verdicts from license- and credential-verified specialists, captures their reasoning including edge cases, and records who authored each example on a signed provenance trail that maps to Annex IV. Data stays in the EU by default.
Best for: teams fine-tuning models in specialised or regulated fields that need gold data whose quality and origin they can prove. See the expert SFT and domain data solution for details.
2. US expert-data marketplaces
These reach real domain experts and scale well. The trade-offs are limited visibility into who authored your data and US data residency, which does not fit EU-regulated fine-tuning.
Best for: US teams that prioritise reach and speed over provenance and EU compliance.
3. BPO and crowd annotation vendors
The large annotation firms are built for volume and cost efficiency. For SFT gold data, the limitation is that a general workforce cannot author authoritative answers in specialised domains, so the ceiling on quality is low where expertise is required.
Best for: high-volume, general-domain SFT data where deep expertise is not needed.
4. Domain-specialist data boutiques
Boutiques focused on a single vertical (for example medical or legal data) bring real subject depth. They can be excellent within their niche, but coverage is narrow, capacity is limited, and provenance and EU fit vary by vendor.
Best for: deep work in exactly the one domain a given boutique specialises in.
5. Synthetic data generation tools
Model-generated SFT data is fast and cheap and helps with coverage. But synthetic gold answers inherit the generating model's errors and blind spots, which is precisely the problem expert data is meant to solve in high-stakes fields.
Best for: augmenting expert data on well-understood tasks, not sourcing the gold standard itself where correctness is hard.
6. In-house SME panels
Employing your own subject-matter experts gives you control and context. It is expensive, slow to scale across domains, and leaves you to build the nuance-capture and provenance layers yourself.
Best for: organisations with a narrow, ongoing domain need and the budget to staff it internally.
7. Academic and freelance expert networks
Hiring credentialed experts directly reaches deep knowledge at a reasonable rate. The overhead is coordination and quality control, with no built-in provenance, so you assemble the reliability layer by hand.
Best for: small, high-expertise gold sets you can manage directly.
How to choose
Anchor on the field and the stakes. For fine-tuning in specialised or regulated domains, author qualifications and provenance are decisive, and synthetic or crowd data will not clear the bar. For general-domain SFT at volume, a crowd vendor or synthetic augmentation may be fine. The test that cuts through the marketing: for any example in your gold set, can the provider name who wrote it and show they were qualified to?