Skip to content

Domain judgment and SFT tasks: expert answers as training data

Pathwize EditorialExpert data and AI evaluation2 min read
GuidesPathwize AI

Sometimes you do not want a score or a comparison, you want the right answer, written by someone qualified to write it. How domain judgment tasks produce gold data.

Preference compares two answers. Evaluation scores one. But sometimes you do not have a good answer to compare or grade at all, you need someone to write it. That is a domain judgment task: an expert produces the ideal answer, or the authoritative verdict, for each item. It is supervised fine-tuning data at its source, the gold response the model will learn from.

Gold data starts with a qualified author

The value of an SFT example is bounded by whoever wrote it. A gold answer in oncology, contract law, or frontier maths is only gold if the author is genuinely qualified to produce it. This is the task type where credentials matter most, because there is no second answer to check against; the expert's output is the standard, so the standard is only as good as the expert.

Two flavours: ideal answer and expert verdict

Some domain judgment tasks ask the expert to write the ideal answer from scratch, the response you wish the model would give. Others ask for a verdict on an existing answer: is this AI-drafted response correct, incorrect, or an edge case, and why? Both produce authoritative expert judgment. The first builds gold responses; the second builds a labelled record of what correct looks like, which is just as useful for training and evaluation.

Why edge cases are worth capturing

The most valuable domain judgments are often the ones that are not clean pass or fail. An expert flagging that an answer is technically correct but misleading, or right in general but wrong for this patient, teaches the model the nuance a binary label would erase. Give experts room to record the edge case rather than forcing a verdict, and capture the reasoning alongside it.

Where domain judgment fits

This is the foundation stage of a data programme. Domain judgment builds the gold set the model is fine-tuned on and the reference answers your evaluations grade against. Preference then tunes behaviour, evaluation measures it, and red teaming stress-tests it, but all of that rests on the quality of the expert answers underneath. Get this layer right and the rest has something solid to stand on.

On Pathwize

A domain judgment task on Pathwize routes each item to credential-verified experts in the relevant field and records the answer or verdict, the reasoning, and the author, as a signed trail. You get gold data whose provenance you can prove, which is exactly what Annex IV expects you to be able to show. Create one from the AI tasks console and set the required expertise.

Frequently asked questions

What is a domain judgment or SFT task?+

A task where a qualified expert writes the ideal answer, or gives the authoritative verdict, for each item. It produces gold responses, the source data for supervised fine-tuning, rather than a score or a comparison.

Why do credentials matter so much for this task type?+

Because there is no second answer to check against. The expert's output is the standard, so the quality of the data is bounded by the qualifications of whoever wrote it.

Should experts be allowed to flag edge cases?+

Yes. The most valuable judgments are often the ones that are not a clean pass or fail. Letting experts record an edge case with their reasoning captures nuance a binary label would erase.

Related reading

See Pathwize on your own data

Source verifiable expert data with provenance built in, EU-native and audit-ready.

Book a demo
← All posts

Related stories

GuidesPathwize AI

How to evaluate an AI data vendor: a due-diligence checklist

The questions that separate a data vendor you can defend to an auditor from one that will cost you a re-labelling project six months in.

Pathwize AIGuides

How to run a human-data pilot before you commit

A pilot is the cheapest way to learn whether a data partner can actually do your hardest work. Here is how to design one that tells you something real.

GuidesPathwize AI

The four types of AI data tasks, explained

Preference, evaluation, red teaming and domain judgment. What each task type does, when to use it, and how they fit together in a data programme.